Model comparison

Gemini 2.0 Pro vs Grok 4.3

Grok 4.3 is the stronger model overall, scoring 43.8 to 39.1 on the Noometry Index.

Last verified . 2 shared benchmarks.

Gemini 2.0 Pro Google

39.1

Rank #173 Confirmed

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Gemini 2.0 Pro scores higher in 1 category and Grok 4.3 in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok 4.3 leads 52.5 to 36.5.
  • The biggest single-benchmark swing is GPQA Diamond: 65.7% for Gemini 2.0 Pro and 88.8% for Grok 4.3.

Side by side

Gemini 2.0 Pro and Grok 4.3 specifications
Gemini 2.0 ProGrok 4.3
ProviderGooglexAI
Noometry Index39.143.8
Released2025-02-052026-04-17
WeightsProprietaryProprietary
Context window—1M
Max output—30K
Input $ / M tokens—$1.25
Output $ / M tokens—$2.50
Results tracked1440

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.3 leads

Gemini 2.0 Pro: 37.8 (#187), Grok 4.3: 41.6 (#121)

Coding benchmarks
BenchmarkGemini 2.0 ProGrok 4.3
Aider Polyglot35.6%—
LMArena WebDev—1357
SciCode—47.3%
WeirdML—49.9%
LiveBench Coding63.5%—
LMArena Coding—1415
ALE-Bench—944.17

Agentic & Tool Use Not comparable

Gemini 2.0 Pro: —, Grok 4.3: 27.7 (#99)

Agentic & Tool Use benchmarks
BenchmarkGemini 2.0 ProGrok 4.3
GDP.pdf—8%
LMArena Search—1165
Vending-Bench 2—35.26

Reasoning Grok 4.3 leads

Gemini 2.0 Pro: 22.3 (#198), Grok 4.3: 35.9 (#68)

Reasoning benchmarks
BenchmarkGemini 2.0 ProGrok 4.3
Epoch Capabilities Index135.06149.16
NYT Connections (extended)—55.2%
CritPt—8%
Chess Puzzles—25%
EnigmaEval0.7%—
LiveBench Reasoning60.1%—
LMArena Hard Prompts—1396
DTBench—90.7%
LiveBench Data Analysis68%—
LMCA—38.3%
ForecastBench—60.3
LiveBench65.1%—

Math Grok 4.3 leads

Gemini 2.0 Pro: 39.7 (#100), Grok 4.3: 46.0 (#74)

Math benchmarks
BenchmarkGemini 2.0 ProGrok 4.3
FrontierMath (Tiers 1-3)—42.8%
FrontierMath Tier 4—14.6%
OTIS Mock AIME 2024-2025—93.3%
ProofBench—11%
LiveBench Math71%—
LMArena Math—1388
MATH Level 583.5%—

Knowledge Grok 4.3 leads

Gemini 2.0 Pro: 36.5 (#167), Grok 4.3: 52.5 (#62)

Knowledge benchmarks
BenchmarkGemini 2.0 ProGrok 4.3
GPQA Diamond65.7%88.8%
SimpleQA Verified—33.2%
Confabulations18.4%—
LMArena Expert—1385

Multimodal Not comparable

Gemini 2.0 Pro: —, Grok 4.3: 31.6 (#104)

Multimodal benchmarks
BenchmarkGemini 2.0 ProGrok 4.3
LMArena Vision—1229
Blueprint-Bench 2—0%

Multilingual Not comparable

Gemini 2.0 Pro: —, Grok 4.3: 50.5 (#120)

Multilingual benchmarks
BenchmarkGemini 2.0 ProGrok 4.3
LMArena Non-English—1385
LMArena Chinese—1422
LMArena French—1412
LMArena German—1395
LMArena Japanese—1379
LMArena Korean—1356
LMArena Russian—1399
LMArena Spanish—1398

Instruction Following Gemini 2.0 Pro leads

Gemini 2.0 Pro: 75.5 (#59), Grok 4.3: 72.1 (#140)

Instruction Following benchmarks
BenchmarkGemini 2.0 ProGrok 4.3
LiveBench Instruction Following83.4%—
LMArena Instruction Following—1366

Long Context Grok 4.3 leads

Gemini 2.0 Pro: 29.2 (#292), Grok 4.3: 42.5 (#123)

Long Context benchmarks
BenchmarkGemini 2.0 ProGrok 4.3
Fiction.LiveBench41.7%—
LMArena Longer Query—1393

Writing & Preference Grok 4.3 leads

Gemini 2.0 Pro: 52.7 (#165), Grok 4.3: 58.5 (#118)

Writing & Preference benchmarks
BenchmarkGemini 2.0 ProGrok 4.3
LMArena Text—1397
LMArena Creative Writing—1380
EQ-Bench 4—1075
LMArena Multi-Turn—1406
LiveBench Language44.9%—

Frequently asked questions

Is Gemini 2.0 Pro better than Grok 4.3?

Grok 4.3 is the stronger model overall, scoring 43.8 to 39.1 on the Noometry Index.

Is Gemini 2.0 Pro or Grok 4.3 better for coding?

Grok 4.3 scores higher on coding benchmarks: 41.6 versus 37.8 in the Noometry coding category.

How many benchmarks do Gemini 2.0 Pro and Grok 4.3 share?

2 benchmarks have published results for both models. Gemini 2.0 Pro has 14 scored results on Noometry and Grok 4.3 has 40.

Related comparisons

Go deeper