Model comparison

Gemini 2.0 Pro vs Grok 4.7

Grok 4.7 is the stronger model overall, scoring 53.1 to 39.1 on the Noometry Index.

Last verified . 2 shared benchmarks.

Gemini 2.0 Pro Google

39.1

Rank #173 Confirmed

Grok 4.7 xAI

53.1

Rank #37 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Gemini 2.0 Pro scores higher in 1 category and Grok 4.7 in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.7 leads 49.1 to 22.3.
  • The biggest single-benchmark swing is GPQA Diamond: 65.7% for Gemini 2.0 Pro and 92.7% for Grok 4.7.

Side by side

Gemini 2.0 Pro and Grok 4.7 specifications
Gemini 2.0 ProGrok 4.7
ProviderGooglexAI
Noometry Index39.153.1
Released2025-02-052026-09-21
WeightsProprietaryProprietary
Context window—500K
Max output—500K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked1439

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.7 leads

Gemini 2.0 Pro: 37.8 (#187), Grok 4.7: 58.0 (#18)

Coding benchmarks
BenchmarkGemini 2.0 ProGrok 4.7
FrontierCode—47.6%
Aider Polyglot35.6%—
CursorBench—46.3%
LMArena WebDev—1639
FrontierSWE—29.5%
SciCode—57.8%
LiveBench Coding63.5%—
LMArena Coding—1427

Agentic & Tool Use Not comparable

Gemini 2.0 Pro: —, Grok 4.7: 36.7 (#37)

Agentic & Tool Use benchmarks
BenchmarkGemini 2.0 ProGrok 4.7
APEX-Agents—54.6%
GDP.pdf—22.8%
Vending-Bench 2—10,537

Reasoning Grok 4.7 leads

Gemini 2.0 Pro: 22.3 (#198), Grok 4.7: 49.1 (#40)

Reasoning benchmarks
BenchmarkGemini 2.0 ProGrok 4.7
Epoch Capabilities Index135.06153.53
NYT Connections (extended)—76.8%
CritPt—18%
Chess Puzzles—38%
EnigmaEval0.7%—
LiveBench Reasoning60.1%—
LMArena Hard Prompts—1413
Mystery Game Puzzles—29%
DTBench—96%
LiveBench Data Analysis68%—
LMCA—49.4%
LiveBench65.1%—

Math Grok 4.7 leads

Gemini 2.0 Pro: 39.7 (#100), Grok 4.7: 57.8 (#39)

Math benchmarks
BenchmarkGemini 2.0 ProGrok 4.7
FrontierMath (Tiers 1-3)—53%
FrontierMath Tier 4—17.1%
OTIS Mock AIME 2024-2025—98.1%
ProofBench—34%
LiveBench Math71%—
LMArena Math—1407
MATH Level 583.5%—

Knowledge Grok 4.7 leads

Gemini 2.0 Pro: 36.5 (#167), Grok 4.7: 62.8 (#22)

Knowledge benchmarks
BenchmarkGemini 2.0 ProGrok 4.7
GPQA Diamond65.7%92.7%
SimpleQA Verified—56%
Confabulations18.4%—
LMArena Expert—1422

Multimodal Not comparable

Gemini 2.0 Pro: —, Grok 4.7: 35.5 (#87)

Multimodal benchmarks
BenchmarkGemini 2.0 ProGrok 4.7
LMArena Vision—1228
Blueprint-Bench 2—32.5%
Furniture Assembly—20.8%

Multilingual Not comparable

Gemini 2.0 Pro: —, Grok 4.7: 50.8 (#116)

Multilingual benchmarks
BenchmarkGemini 2.0 ProGrok 4.7
LMArena Non-English—1389
LMArena Chinese—1455
LMArena French—1455
LMArena Russian—1397
LMArena Spanish—1400

Instruction Following Gemini 2.0 Pro leads

Gemini 2.0 Pro: 75.5 (#59), Grok 4.7: 74.1 (#105)

Instruction Following benchmarks
BenchmarkGemini 2.0 ProGrok 4.7
LiveBench Instruction Following83.4%—
LMArena Instruction Following—1404

Long Context Grok 4.7 leads

Gemini 2.0 Pro: 29.2 (#292), Grok 4.7: 43.1 (#104)

Long Context benchmarks
BenchmarkGemini 2.0 ProGrok 4.7
Fiction.LiveBench41.7%—
LMArena Longer Query—1413

Writing & Preference Grok 4.7 leads

Gemini 2.0 Pro: 52.7 (#165), Grok 4.7: 70.0 (#24)

Writing & Preference benchmarks
BenchmarkGemini 2.0 ProGrok 4.7
LMArena Text—1399
LMArena Creative Writing—1391
EQ-Bench Creative Writing—2007
LMArena Multi-Turn—1393
LiveBench Language44.9%—

Frequently asked questions

Is Gemini 2.0 Pro better than Grok 4.7?

Grok 4.7 is the stronger model overall, scoring 53.1 to 39.1 on the Noometry Index.

Is Gemini 2.0 Pro or Grok 4.7 better for coding?

Grok 4.7 scores higher on coding benchmarks: 58.0 versus 37.8 in the Noometry coding category.

How many benchmarks do Gemini 2.0 Pro and Grok 4.7 share?

2 benchmarks have published results for both models. Gemini 2.0 Pro has 14 scored results on Noometry and Grok 4.7 has 39.

Related comparisons

Go deeper