Model comparison

Gemini 2.5 Flash vs Grok Build 0.1

Gemini 2.5 Flash is the stronger model overall, scoring 39.3 to 36.4 on the Noometry Index.

Last verified . 1 shared benchmarks.

Gemini 2.5 Flash Google

39.3

Rank #170 Confirmed

Grok Build 0.1 xAI

36.4

Rank #216 Reported

Summary

  • They share 1 benchmark with published results for both. Gemini 2.5 Flash scores higher in 1 category and Grok Build 0.1 in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok Build 0.1 leads 32.2 to 18.1.
  • The biggest single-benchmark swing is CritPt: 1.1% for Gemini 2.5 Flash and 9.1% for Grok Build 0.1.
  • Gemini 2.5 Flash is cheaper at $0.30 / $2.50 per million input/output tokens, against $1 / $2 for Grok Build 0.1.
  • Gemini 2.5 Flash accepts more context: 1.05M tokens versus 256K.

Side by side

Gemini 2.5 Flash and Grok Build 0.1 specifications
Gemini 2.5 FlashGrok Build 0.1
ProviderGooglexAI
Noometry Index39.336.4
Released2025-04-172026-04-16
WeightsProprietaryProprietary
Context window1.05M256K
Max output66K256K
Input $ / M tokens$0.30$1
Output $ / M tokens$2.50$2
Results tracked543

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok Build 0.1 leads

Gemini 2.5 Flash: 35.8 (#220), Grok Build 0.1: 43.1 (#91)

Coding benchmarks
BenchmarkGemini 2.5 FlashGrok Build 0.1
SWE-bench Verified (bash only)28.7%—
Aider Polyglot55.1%—
SciCode—50.2%
WeirdML41.9%—
LMArena Coding1424—
ALE-Bench661.88—

Agentic & Tool Use Gemini 2.5 Flash leads

Gemini 2.5 Flash: 30.8 (#74), Grok Build 0.1: 22.7 (#129)

Agentic & Tool Use benchmarks
BenchmarkGemini 2.5 FlashGrok Build 0.1
Terminal-Bench17.1%—
Berkeley Function Calling Leaderboard56.2%—
TheAgentCompany41.1%—
BALROG33.5%—
GBAEval—2.4%
Vending-Bench 2548.84—

Reasoning Grok Build 0.1 leads

Gemini 2.5 Flash: 18.1 (#286), Grok Build 0.1: 32.2 (#77)

Reasoning benchmarks
BenchmarkGemini 2.5 FlashGrok Build 0.1
CritPt1.1%9.1%
ARC-AGI-22.5%—
SimpleBench41.2%—
Kagi LLM Benchmark56.8%—
ARC-AGI-133.3%—
EnigmaEval2.7%—
LMArena Hard Prompts1422—
DTBench76.5%—
LMCA27.5%—
Epoch Capabilities Index143.03—
ForecastBench60.6—

Math Not comparable

Gemini 2.5 Flash: 39.9 (#98), Grok Build 0.1: —

Math benchmarks
BenchmarkGemini 2.5 FlashGrok Build 0.1
OTIS Mock AIME 2024-202573.1%—
Omni-MATH38.5%—
LMArena Math1415—
FrontierMath (Feb 2025 set)4.8%—
FrontierMath Tier 4 (v1)4.2%—

Knowledge Not comparable

Gemini 2.5 Flash: 36.4 (#168), Grok Build 0.1: —

Knowledge benchmarks
BenchmarkGemini 2.5 FlashGrok Build 0.1
Humanity's Last Exam12.1%—
MMLU-Pro63.9%—
Confabulations16.8%—
Vectara Hallucination Rate7.8%—
GPQA (HELM)39%—
LMArena Expert1426—

Multimodal Not comparable

Gemini 2.5 Flash: 41.8 (#32), Grok Build 0.1: —

Multimodal benchmarks
BenchmarkGemini 2.5 FlashGrok Build 0.1
LMArena Vision1253—
GeoBench76%—
VPCT46.2%—
SpatialViz-Bench36.9%—

Multilingual Not comparable

Gemini 2.5 Flash: 52.3 (#88), Grok Build 0.1: —

Multilingual benchmarks
BenchmarkGemini 2.5 FlashGrok Build 0.1
LMArena Non-English1409—
LMArena Chinese1450—
LMArena French1433—
LMArena German1418—
LMArena Japanese1405—
LMArena Korean1385—
LMArena Russian1415—
LMArena Spanish1421—

Instruction Following Not comparable

Gemini 2.5 Flash: 75.7 (#54), Grok Build 0.1: —

Instruction Following benchmarks
BenchmarkGemini 2.5 FlashGrok Build 0.1
IFEval89.8%—
LMArena Instruction Following1405—

Long Context Not comparable

Gemini 2.5 Flash: 47.5 (#17), Grok Build 0.1: —

Long Context benchmarks
BenchmarkGemini 2.5 FlashGrok Build 0.1
Fiction.LiveBench77.8%—
LMArena Longer Query1419—

Writing & Preference Not comparable

Gemini 2.5 Flash: 53.8 (#157), Grok Build 0.1: —

Writing & Preference benchmarks
BenchmarkGemini 2.5 FlashGrok Build 0.1
LMArena Text1417—
LMArena Creative Writing1400—
Short-Story Creative Writing76.5%—
EQ-Bench Creative Writing1137—
WildBench81.7%—
LMArena Multi-Turn1408—

Frequently asked questions

Is Gemini 2.5 Flash better than Grok Build 0.1?

Gemini 2.5 Flash is the stronger model overall, scoring 39.3 to 36.4 on the Noometry Index.

Which is cheaper, Gemini 2.5 Flash or Grok Build 0.1?

Gemini 2.5 Flash is cheaper. It lists at $0.30 per million input tokens and $2.50 per million output tokens; Grok Build 0.1 lists at $1 and $2.

Is Gemini 2.5 Flash or Grok Build 0.1 better for coding?

Grok Build 0.1 scores higher on coding benchmarks: 43.1 versus 35.8 in the Noometry coding category.

Which has the bigger context window?

Gemini 2.5 Flash does, with 1.05M tokens against 256K.

How many benchmarks do Gemini 2.5 Flash and Grok Build 0.1 share?

1 benchmark has published results for both models. Gemini 2.5 Flash has 54 scored results on Noometry and Grok Build 0.1 has 3.

Related comparisons

Go deeper