Model comparison

Gemini 1.0 Pro vs Grok 4.5

Grok 4.5 is the stronger model overall, scoring 55.0 to 27.3 on the Noometry Index.

Last verified . 20 shared benchmarks.

Gemini 1.0 Pro Google

27.3

Rank #332 Confirmed

Grok 4.5 xAI

55.0

Rank #25 Confirmed

Summary

  • They share 20 benchmarks with published results for both. Gemini 1.0 Pro scores higher in 0 categories and Grok 4.5 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Grok 4.5 leads 60.9 to 9.3.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 1.1% for Gemini 1.0 Pro and 97.8% for Grok 4.5.

Side by side

Gemini 1.0 Pro and Grok 4.5 specifications
Gemini 1.0 ProGrok 4.5
ProviderGooglexAI
Noometry Index27.355.0
Released2023-12-132026-07-08
WeightsProprietaryProprietary
Context window—500K
Max output—500K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked2452

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.5 leads

Gemini 1.0 Pro: 32.2 (#275), Grok 4.5: 52.2 (#35)

Coding benchmarks
BenchmarkGemini 1.0 ProGrok 4.5
LMArena Coding11081474
DeepSWE—53.8%
FrontierCode—42.4%
LMArena WebDev—1553
SciCode—54.1%
WeirdML—46.4%
ALE-Bench—1,309
HumanEval+55.5%—
MBPP+61.4%—

Agentic & Tool Use Not comparable

Gemini 1.0 Pro: —, Grok 4.5: 44.4 (#17)

Agentic & Tool Use benchmarks
BenchmarkGemini 1.0 ProGrok 4.5
APEX-Agents—56.2%
τ²-bench Banking—47.9%
PostTrainBench—23.4%
GBAEval—65.4%
GDP.pdf—14%
LMArena Search—1213
Vending-Bench 2—3,887

Reasoning Grok 4.5 leads

Gemini 1.0 Pro: 17.1 (#296), Grok 4.5: 56.1 (#25)

Reasoning benchmarks
BenchmarkGemini 1.0 ProGrok 4.5
LMArena Hard Prompts11091462
DTBench45.9%96.5%
Epoch Capabilities Index117.04153.92
ARC-AGI-2—52.6%
SimpleBench—70%
Kagi LLM Benchmark—83.5%
NYT Connections (extended)—79.9%
ARC-AGI-1—87.2%
CritPt—15.4%
Chess Puzzles—36%
LMCA—45.2%
Surface Evolver Bench—74.4%

Math Grok 4.5 leads

Gemini 1.0 Pro: 9.3 (#321), Grok 4.5: 60.9 (#35)

Math benchmarks
BenchmarkGemini 1.0 ProGrok 4.5
OTIS Mock AIME 2024-20251.1%97.8%
LMArena Math11321459
FrontierMath (Tiers 1-3)—57.2%
FrontierMath Tier 4—24.4%
ProofBench—31%
MATH Level 511.2%—

Knowledge Grok 4.5 leads

Gemini 1.0 Pro: 15.6 (#291), Grok 4.5: 62.3 (#24)

Knowledge benchmarks
BenchmarkGemini 1.0 ProGrok 4.5
GPQA Diamond34%93.4%
LMArena Expert10591466
SimpleQA Verified—48.3%
MMLU70%—

Multimodal Not comparable

Gemini 1.0 Pro: —, Grok 4.5: 37.6 (#72)

Multimodal benchmarks
BenchmarkGemini 1.0 ProGrok 4.5
LMArena Vision—1288
Blueprint-Bench 2—27.3%
Furniture Assembly—22.5%
LMArena Document—1452

Multilingual Grok 4.5 leads

Gemini 1.0 Pro: 33.4 (#252), Grok 4.5: 54.4 (#42)

Multilingual benchmarks
BenchmarkGemini 1.0 ProGrok 4.5
LMArena Non-English11381440
LMArena Chinese11241496
LMArena French11451456
LMArena German11251446
LMArena Japanese10231428
LMArena Russian11861448
LMArena Spanish11191450
LMArena Korean—1404

Instruction Following Grok 4.5 leads

Gemini 1.0 Pro: 57.6 (#267), Grok 4.5: 76.0 (#48)

Instruction Following benchmarks
BenchmarkGemini 1.0 ProGrok 4.5
LMArena Instruction Following11141446

Long Context Grok 4.5 leads

Gemini 1.0 Pro: 34.3 (#249), Grok 4.5: 44.8 (#56)

Long Context benchmarks
BenchmarkGemini 1.0 ProGrok 4.5
LMArena Longer Query11321463

Writing & Preference Grok 4.5 leads

Gemini 1.0 Pro: 36.0 (#264), Grok 4.5: 65.8 (#42)

Writing & Preference benchmarks
BenchmarkGemini 1.0 ProGrok 4.5
LMArena Text11491448
LMArena Creative Writing11311442
LMArena Multi-Turn11391456
EQ-Bench Creative Writing—1579

Frequently asked questions

Is Gemini 1.0 Pro better than Grok 4.5?

Grok 4.5 is the stronger model overall, scoring 55.0 to 27.3 on the Noometry Index.

Is Gemini 1.0 Pro or Grok 4.5 better for coding?

Grok 4.5 scores higher on coding benchmarks: 52.2 versus 32.2 in the Noometry coding category.

How many benchmarks do Gemini 1.0 Pro and Grok 4.5 share?

20 benchmarks have published results for both models. Gemini 1.0 Pro has 24 scored results on Noometry and Grok 4.5 has 52.

Related comparisons

Go deeper