Model comparison

Gemini 2.0 Pro vs Grok 4.1

Grok 4.1 is the stronger model overall, scoring 41.5 to 39.1 on the Noometry Index.

Last verified . 0 shared benchmarks.

Gemini 2.0 Pro Google

39.1

Rank #173 Confirmed

Grok 4.1 xAI

41.5

Rank #134 Confirmed

Summary

  • The widest gap is in long context, where Grok 4.1 leads 43.2 to 29.2.

Side by side

Gemini 2.0 Pro and Grok 4.1 specifications
Gemini 2.0 ProGrok 4.1
ProviderGooglexAI
Noometry Index39.141.5
Released2025-02-052025-11-17
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1419

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 2.0 Pro leads

Gemini 2.0 Pro: 37.8 (#187), Grok 4.1: 33.7 (#253)

Coding benchmarks
BenchmarkGemini 2.0 ProGrok 4.1
Aider Polyglot35.6%—
LMArena WebDev—1214
LiveBench Coding63.5%—
LMArena Coding—1445

Agentic & Tool Use Not comparable

Gemini 2.0 Pro: —, Grok 4.1: 34.1 (#49)

Agentic & Tool Use benchmarks
BenchmarkGemini 2.0 ProGrok 4.1
Cybench—39%

Reasoning Grok 4.1 leads

Gemini 2.0 Pro: 22.3 (#198), Grok 4.1: 29.5 (#91)

Reasoning benchmarks
BenchmarkGemini 2.0 ProGrok 4.1
EnigmaEval0.7%—
LiveBench Reasoning60.1%—
LMArena Hard Prompts—1435
LiveBench Data Analysis68%—
Epoch Capabilities Index135.06—
LiveBench65.1%—

Math Too close to call

Gemini 2.0 Pro: 39.7 (#100), Grok 4.1: 38.9 (#120)

Math benchmarks
BenchmarkGemini 2.0 ProGrok 4.1
LiveBench Math71%—
LMArena Math—1422
MATH Level 583.5%—

Knowledge Grok 4.1 leads

Gemini 2.0 Pro: 36.5 (#167), Grok 4.1: 39.5 (#133)

Knowledge benchmarks
BenchmarkGemini 2.0 ProGrok 4.1
GPQA Diamond65.7%—
Confabulations18.4%—
LMArena Expert—1417

Multilingual Not comparable

Gemini 2.0 Pro: —, Grok 4.1: 53.4 (#68)

Multilingual benchmarks
BenchmarkGemini 2.0 ProGrok 4.1
LMArena Non-English—1425
LMArena Chinese—1465
LMArena French—1448
LMArena German—1446
LMArena Japanese—1397
LMArena Korean—1407
LMArena Russian—1434
LMArena Spanish—1438

Instruction Following Gemini 2.0 Pro leads

Gemini 2.0 Pro: 75.5 (#59), Grok 4.1: 73.8 (#111)

Instruction Following benchmarks
BenchmarkGemini 2.0 ProGrok 4.1
LiveBench Instruction Following83.4%—
LMArena Instruction Following—1400

Long Context Grok 4.1 leads

Gemini 2.0 Pro: 29.2 (#292), Grok 4.1: 43.2 (#100)

Long Context benchmarks
BenchmarkGemini 2.0 ProGrok 4.1
Fiction.LiveBench41.7%—
LMArena Longer Query—1416

Writing & Preference Grok 4.1 leads

Gemini 2.0 Pro: 52.7 (#165), Grok 4.1: 62.4 (#75)

Writing & Preference benchmarks
BenchmarkGemini 2.0 ProGrok 4.1
LMArena Text—1437
LMArena Creative Writing—1411
LMArena Multi-Turn—1437
LiveBench Language44.9%—

Frequently asked questions

Is Gemini 2.0 Pro better than Grok 4.1?

Grok 4.1 is the stronger model overall, scoring 41.5 to 39.1 on the Noometry Index.

Is Gemini 2.0 Pro or Grok 4.1 better for coding?

Gemini 2.0 Pro scores higher on coding benchmarks: 37.8 versus 33.7 in the Noometry coding category.

How many benchmarks do Gemini 2.0 Pro and Grok 4.1 share?

0 benchmarks have published results for both models. Gemini 2.0 Pro has 14 scored results on Noometry and Grok 4.1 has 19.

Related comparisons

Go deeper