Model comparison

Gemini 2.0 Pro vs Grok 4.5

Grok 4.5 is the stronger model overall, scoring 55.0 to 39.1 on the Noometry Index.

Last verified . 2 shared benchmarks.

Gemini 2.0 Pro Google

39.1

Rank #173 Confirmed

Grok 4.5 xAI

55.0

Rank #25 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Gemini 2.0 Pro scores higher in 0 categories and Grok 4.5 in 7 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.5 leads 56.1 to 22.3.
  • The biggest single-benchmark swing is GPQA Diamond: 65.7% for Gemini 2.0 Pro and 93.4% for Grok 4.5.

Side by side

Gemini 2.0 Pro and Grok 4.5 specifications
Gemini 2.0 ProGrok 4.5
ProviderGooglexAI
Noometry Index39.155.0
Released2025-02-052026-07-08
WeightsProprietaryProprietary
Context window—500K
Max output—500K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked1452

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.5 leads

Gemini 2.0 Pro: 37.8 (#187), Grok 4.5: 52.2 (#35)

Coding benchmarks
BenchmarkGemini 2.0 ProGrok 4.5
DeepSWE—53.8%
FrontierCode—42.4%
Aider Polyglot35.6%—
LMArena WebDev—1553
SciCode—54.1%
WeirdML—46.4%
LiveBench Coding63.5%—
LMArena Coding—1474
ALE-Bench—1,309

Agentic & Tool Use Not comparable

Gemini 2.0 Pro: —, Grok 4.5: 44.4 (#17)

Agentic & Tool Use benchmarks
BenchmarkGemini 2.0 ProGrok 4.5
APEX-Agents—56.2%
τ²-bench Banking—47.9%
PostTrainBench—23.4%
GBAEval—65.4%
GDP.pdf—14%
LMArena Search—1213
Vending-Bench 2—3,887

Reasoning Grok 4.5 leads

Gemini 2.0 Pro: 22.3 (#198), Grok 4.5: 56.1 (#25)

Reasoning benchmarks
BenchmarkGemini 2.0 ProGrok 4.5
Epoch Capabilities Index135.06153.92
ARC-AGI-2—52.6%
SimpleBench—70%
Kagi LLM Benchmark—83.5%
NYT Connections (extended)—79.9%
ARC-AGI-1—87.2%
CritPt—15.4%
Chess Puzzles—36%
EnigmaEval0.7%—
LiveBench Reasoning60.1%—
LMArena Hard Prompts—1462
DTBench—96.5%
LiveBench Data Analysis68%—
LMCA—45.2%
Surface Evolver Bench—74.4%
LiveBench65.1%—

Math Grok 4.5 leads

Gemini 2.0 Pro: 39.7 (#100), Grok 4.5: 60.9 (#35)

Math benchmarks
BenchmarkGemini 2.0 ProGrok 4.5
FrontierMath (Tiers 1-3)—57.2%
FrontierMath Tier 4—24.4%
OTIS Mock AIME 2024-2025—97.8%
ProofBench—31%
LiveBench Math71%—
LMArena Math—1459
MATH Level 583.5%—

Knowledge Grok 4.5 leads

Gemini 2.0 Pro: 36.5 (#167), Grok 4.5: 62.3 (#24)

Knowledge benchmarks
BenchmarkGemini 2.0 ProGrok 4.5
GPQA Diamond65.7%93.4%
SimpleQA Verified—48.3%
Confabulations18.4%—
LMArena Expert—1466

Multimodal Not comparable

Gemini 2.0 Pro: —, Grok 4.5: 37.6 (#72)

Multimodal benchmarks
BenchmarkGemini 2.0 ProGrok 4.5
LMArena Vision—1288
Blueprint-Bench 2—27.3%
Furniture Assembly—22.5%
LMArena Document—1452

Multilingual Not comparable

Gemini 2.0 Pro: —, Grok 4.5: 54.4 (#42)

Multilingual benchmarks
BenchmarkGemini 2.0 ProGrok 4.5
LMArena Non-English—1440
LMArena Chinese—1496
LMArena French—1456
LMArena German—1446
LMArena Japanese—1428
LMArena Korean—1404
LMArena Russian—1448
LMArena Spanish—1450

Instruction Following Too close to call

Gemini 2.0 Pro: 75.5 (#59), Grok 4.5: 76.0 (#48)

Instruction Following benchmarks
BenchmarkGemini 2.0 ProGrok 4.5
LiveBench Instruction Following83.4%—
LMArena Instruction Following—1446

Long Context Grok 4.5 leads

Gemini 2.0 Pro: 29.2 (#292), Grok 4.5: 44.8 (#56)

Long Context benchmarks
BenchmarkGemini 2.0 ProGrok 4.5
Fiction.LiveBench41.7%—
LMArena Longer Query—1463

Writing & Preference Grok 4.5 leads

Gemini 2.0 Pro: 52.7 (#165), Grok 4.5: 65.8 (#42)

Writing & Preference benchmarks
BenchmarkGemini 2.0 ProGrok 4.5
LMArena Text—1448
LMArena Creative Writing—1442
EQ-Bench Creative Writing—1579
LMArena Multi-Turn—1456
LiveBench Language44.9%—

Frequently asked questions

Is Gemini 2.0 Pro better than Grok 4.5?

Grok 4.5 is the stronger model overall, scoring 55.0 to 39.1 on the Noometry Index.

Is Gemini 2.0 Pro or Grok 4.5 better for coding?

Grok 4.5 scores higher on coding benchmarks: 52.2 versus 37.8 in the Noometry coding category.

How many benchmarks do Gemini 2.0 Pro and Grok 4.5 share?

2 benchmarks have published results for both models. Gemini 2.0 Pro has 14 scored results on Noometry and Grok 4.5 has 52.

Related comparisons

Go deeper