Model comparison

Gemini 1.5 Flash 8B vs Grok 4.1 Fast

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 29.9 on the Noometry Index.

Last verified . 19 shared benchmarks.

Gemini 1.5 Flash 8B Google

29.9

Rank #301 Confirmed

Grok 4.1 Fast xAI

41.4

Rank #136 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Gemini 1.5 Flash 8B scores higher in 1 category and Grok 4.1 Fast in 8 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.1 Fast leads 43.4 to 20.0.
  • The biggest single-benchmark swing is DTBench: 50% for Gemini 1.5 Flash 8B and 87.7% for Grok 4.1 Fast.

Side by side

Gemini 1.5 Flash 8B and Grok 4.1 Fast specifications
Gemini 1.5 Flash 8BGrok 4.1 Fast
ProviderGooglexAI
Noometry Index29.941.4
Released2024-10-032025-06-27
WeightsProprietaryProprietary
Context window—128K
Max output—30K
Input $ / M tokens—$0.20
Output $ / M tokens—$0.50
Results tracked2132

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 1.5 Flash 8B leads

Gemini 1.5 Flash 8B: 35.5 (#225), Grok 4.1 Fast: 34.1 (#245)

Coding benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.1 Fast
LMArena Coding12181411
LMArena WebDev—1242
ALE-Bench—394.93

Agentic & Tool Use Not comparable

Gemini 1.5 Flash 8B: —, Grok 4.1 Fast: 36.3 (#39)

Agentic & Tool Use benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.1 Fast
Berkeley Function Calling Leaderboard—69.6%
τ²-bench Banking—13.1%
LMArena Search—1171
Vending-Bench 2—1,107

Reasoning Grok 4.1 Fast leads

Gemini 1.5 Flash 8B: 20.0 (#244), Grok 4.1 Fast: 43.4 (#49)

Reasoning benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.1 Fast
LMArena Hard Prompts12091407
DTBench50%87.7%
SimpleBench—56%
NYT Connections (extended)—87.4%
ForecastBench—61

Math Grok 4.1 Fast leads

Gemini 1.5 Flash 8B: 14.2 (#302), Grok 4.1 Fast: 31.9 (#221)

Math benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.1 Fast
LMArena Math12071408
MathArena Final-Answer Competitions—60.9%
OTIS Mock AIME 2024-20254.6%—
ProofBench—4%

Knowledge Grok 4.1 Fast leads

Gemini 1.5 Flash 8B: 16.0 (#289), Grok 4.1 Fast: 33.1 (#207)

Knowledge benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.1 Fast
LMArena Expert11851399
GPQA Diamond33%—
Vectara Hallucination Rate—17.8%

Multimodal Grok 4.1 Fast leads

Gemini 1.5 Flash 8B: 28.2 (#115), Grok 4.1 Fast: 37.0 (#76)

Multimodal benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.1 Fast
LMArena Vision10441201

Multilingual Grok 4.1 Fast leads

Gemini 1.5 Flash 8B: 38.5 (#229), Grok 4.1 Fast: 51.0 (#114)

Multilingual benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.1 Fast
LMArena Non-English12151391
LMArena Chinese12311441
LMArena French12341415
LMArena German12061404
LMArena Japanese11501349
LMArena Korean11401361
LMArena Russian12361387
LMArena Spanish12121413

Instruction Following Grok 4.1 Fast leads

Gemini 1.5 Flash 8B: 62.8 (#236), Grok 4.1 Fast: 72.7 (#133)

Instruction Following benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.1 Fast
LMArena Instruction Following11991376

Long Context Grok 4.1 Fast leads

Gemini 1.5 Flash 8B: 37.0 (#225), Grok 4.1 Fast: 42.4 (#126)

Long Context benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.1 Fast
LMArena Longer Query12191390

Writing & Preference Grok 4.1 Fast leads

Gemini 1.5 Flash 8B: 42.8 (#232), Grok 4.1 Fast: 57.2 (#131)

Writing & Preference benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.1 Fast
LMArena Text12261408
LMArena Creative Writing12181394
LMArena Multi-Turn11851389
EQ-Bench Creative Writing—1327

Frequently asked questions

Is Gemini 1.5 Flash 8B better than Grok 4.1 Fast?

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 29.9 on the Noometry Index.

Is Gemini 1.5 Flash 8B or Grok 4.1 Fast better for coding?

Gemini 1.5 Flash 8B scores higher on coding benchmarks: 35.5 versus 34.1 in the Noometry coding category.

How many benchmarks do Gemini 1.5 Flash 8B and Grok 4.1 Fast share?

19 benchmarks have published results for both models. Gemini 1.5 Flash 8B has 21 scored results on Noometry and Grok 4.1 Fast has 32.

Related comparisons

Go deeper