Model comparison

Gemini 1.5 Flash 8B vs Grok 4.7

Grok 4.7 is the stronger model overall, scoring 53.1 to 29.9 on the Noometry Index.

Last verified . 18 shared benchmarks.

Gemini 1.5 Flash 8B Google

29.9

Rank #301 Confirmed

Grok 4.7 xAI

53.1

Rank #37 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Gemini 1.5 Flash 8B scores higher in 0 categories and Grok 4.7 in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok 4.7 leads 62.8 to 16.0.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 4.6% for Gemini 1.5 Flash 8B and 98.1% for Grok 4.7.

Side by side

Gemini 1.5 Flash 8B and Grok 4.7 specifications
Gemini 1.5 Flash 8BGrok 4.7
ProviderGooglexAI
Noometry Index29.953.1
Released2024-10-032026-09-21
WeightsProprietaryProprietary
Context window—500K
Max output—500K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked2139

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.7 leads

Gemini 1.5 Flash 8B: 35.5 (#225), Grok 4.7: 58.0 (#18)

Coding benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.7
LMArena Coding12181427
FrontierCode—47.6%
CursorBench—46.3%
LMArena WebDev—1639
FrontierSWE—29.5%
SciCode—57.8%

Agentic & Tool Use Not comparable

Gemini 1.5 Flash 8B: —, Grok 4.7: 36.7 (#37)

Agentic & Tool Use benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.7
APEX-Agents—54.6%
GDP.pdf—22.8%
Vending-Bench 2—10,537

Reasoning Grok 4.7 leads

Gemini 1.5 Flash 8B: 20.0 (#244), Grok 4.7: 49.1 (#40)

Reasoning benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.7
LMArena Hard Prompts12091413
DTBench50%96%
NYT Connections (extended)—76.8%
CritPt—18%
Chess Puzzles—38%
Mystery Game Puzzles—29%
LMCA—49.4%
Epoch Capabilities Index—153.53

Math Grok 4.7 leads

Gemini 1.5 Flash 8B: 14.2 (#302), Grok 4.7: 57.8 (#39)

Math benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.7
OTIS Mock AIME 2024-20254.6%98.1%
LMArena Math12071407
FrontierMath (Tiers 1-3)—53%
FrontierMath Tier 4—17.1%
ProofBench—34%

Knowledge Grok 4.7 leads

Gemini 1.5 Flash 8B: 16.0 (#289), Grok 4.7: 62.8 (#22)

Knowledge benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.7
GPQA Diamond33%92.7%
LMArena Expert11851422
SimpleQA Verified—56%

Multimodal Grok 4.7 leads

Gemini 1.5 Flash 8B: 28.2 (#115), Grok 4.7: 35.5 (#87)

Multimodal benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.7
LMArena Vision10441228
Blueprint-Bench 2—32.5%
Furniture Assembly—20.8%

Multilingual Grok 4.7 leads

Gemini 1.5 Flash 8B: 38.5 (#229), Grok 4.7: 50.8 (#116)

Multilingual benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.7
LMArena Non-English12151389
LMArena Chinese12311455
LMArena French12341455
LMArena Russian12361397
LMArena Spanish12121400
LMArena German1206—
LMArena Japanese1150—
LMArena Korean1140—

Instruction Following Grok 4.7 leads

Gemini 1.5 Flash 8B: 62.8 (#236), Grok 4.7: 74.1 (#105)

Instruction Following benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.7
LMArena Instruction Following11991404

Long Context Grok 4.7 leads

Gemini 1.5 Flash 8B: 37.0 (#225), Grok 4.7: 43.1 (#104)

Long Context benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.7
LMArena Longer Query12191413

Writing & Preference Grok 4.7 leads

Gemini 1.5 Flash 8B: 42.8 (#232), Grok 4.7: 70.0 (#24)

Writing & Preference benchmarks
BenchmarkGemini 1.5 Flash 8BGrok 4.7
LMArena Text12261399
LMArena Creative Writing12181391
LMArena Multi-Turn11851393
EQ-Bench Creative Writing—2007

Frequently asked questions

Is Gemini 1.5 Flash 8B better than Grok 4.7?

Grok 4.7 is the stronger model overall, scoring 53.1 to 29.9 on the Noometry Index.

Is Gemini 1.5 Flash 8B or Grok 4.7 better for coding?

Grok 4.7 scores higher on coding benchmarks: 58.0 versus 35.5 in the Noometry coding category.

How many benchmarks do Gemini 1.5 Flash 8B and Grok 4.7 share?

18 benchmarks have published results for both models. Gemini 1.5 Flash 8B has 21 scored results on Noometry and Grok 4.7 has 39.

Related comparisons

Go deeper