Model comparison

Gemini 2.0 Pro vs o3

o3 is the stronger model overall, scoring 47.5 to 39.1 on the Noometry Index.

Last verified . 7 shared benchmarks.

Gemini 2.0 Pro Google

39.1

Rank #173 Confirmed

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 7 benchmarks with published results for both. Gemini 2.0 Pro scores higher in 1 category and o3 in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in long context, where o3 leads 53.3 to 29.2.
  • The biggest single-benchmark swing is Fiction.LiveBench: 41.7% for Gemini 2.0 Pro and 88.9% for o3.

Side by side

Gemini 2.0 Pro and o3 specifications
Gemini 2.0 Proo3
ProviderGoogleOpenAI
Noometry Index39.147.5
Released2025-02-052025-04-16
WeightsProprietaryProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$2
Output $ / M tokens—$8
Results tracked1463

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3 leads

Gemini 2.0 Pro: 37.8 (#187), o3: 46.8 (#64)

Coding benchmarks
BenchmarkGemini 2.0 Proo3
Aider Polyglot35.6%81.3%
SWE-bench Verified—62.3%
SWE-bench Verified (bash only)—58.4%
GSO—8.8%
WeirdML—52.4%
LiveBench Coding63.5%—
LMArena Coding—1408
CadEval—74%
ALE-Bench—933.55

Agentic & Tool Use Not comparable

Gemini 2.0 Pro: —, o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkGemini 2.0 Proo3
Berkeley Function Calling Leaderboard—63%
GDPval—30.8%
DeepResearch Bench—45.2%
OSWorld—23%
LMArena Search—1144
METR Time Horizons—65.4%

Reasoning o3 leads

Gemini 2.0 Pro: 22.3 (#198), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkGemini 2.0 Proo3
EnigmaEval0.7%13.1%
Epoch Capabilities Index135.06146.86
ARC-AGI-2—6.5%
SimpleBench—53.1%
Kagi LLM Benchmark—67.6%
ARC-AGI-1—60.8%
CritPt—1.4%
Chess Puzzles—38%
LiveBench Reasoning60.1%—
LMArena Hard Prompts—1402
Mystery Game Puzzles—29%
DTBench—84.8%
LiveBench Data Analysis68%—
LMCA—39.7%
ForecastBench—62.5
LiveBench65.1%—

Math o3 leads

Gemini 2.0 Pro: 39.7 (#100), o3: 50.2 (#58)

Math benchmarks
BenchmarkGemini 2.0 Proo3
MATH Level 583.5%97.8%
FrontierMath (Tiers 1-3)—33.3%
OTIS Mock AIME 2024-2025—84.4%
Omni-MATH—71.4%
LiveBench Math71%—
LMArena Math—1426
FrontierMath (Feb 2025 set)—18.7%
FrontierMath Tier 4 (v1)—2.1%

Knowledge o3 leads

Gemini 2.0 Pro: 36.5 (#167), o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkGemini 2.0 Proo3
GPQA Diamond65.7%81.8%
Confabulations18.4%14.4%
Humanity's Last Exam—20.3%
SimpleQA Verified—49.4%
MMLU-Pro—85.9%
GPQA (HELM)—75.3%
LMArena Expert—1402

Multimodal Not comparable

Gemini 2.0 Pro: —, o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkGemini 2.0 Proo3
LMArena Vision—1214
GeoBench—74%
VPCT—52%

Multilingual Not comparable

Gemini 2.0 Pro: —, o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkGemini 2.0 Proo3
LMArena Non-English—1401
LMArena Chinese—1437
LMArena French—1430
LMArena German—1420
LMArena Japanese—1403
LMArena Korean—1370
LMArena Russian—1406
LMArena Spanish—1395

Instruction Following Gemini 2.0 Pro leads

Gemini 2.0 Pro: 75.5 (#59), o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkGemini 2.0 Proo3
LiveBench Instruction Following83.4%—
IFEval—86.9%
LMArena Instruction Following—1368

Long Context o3 leads

Gemini 2.0 Pro: 29.2 (#292), o3: 53.3 (#6)

Long Context benchmarks
BenchmarkGemini 2.0 Proo3
Fiction.LiveBench41.7%88.9%
CL-bench—17.8%
LMArena Longer Query—1372

Writing & Preference o3 leads

Gemini 2.0 Pro: 52.7 (#165), o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkGemini 2.0 Proo3
LMArena Text—1410
LMArena Creative Writing—1359
Short-Story Creative Writing—83.9%
EQ-Bench Creative Writing—1676
WildBench—86.1%
LMArena Multi-Turn—1405
LiveBench Language44.9%—

Frequently asked questions

Is Gemini 2.0 Pro better than o3?

o3 is the stronger model overall, scoring 47.5 to 39.1 on the Noometry Index.

Is Gemini 2.0 Pro or o3 better for coding?

o3 scores higher on coding benchmarks: 46.8 versus 37.8 in the Noometry coding category.

How many benchmarks do Gemini 2.0 Pro and o3 share?

7 benchmarks have published results for both models. Gemini 2.0 Pro has 14 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper