Model comparison

Gemini 1.0 Pro vs o3

o3 is the stronger model overall, scoring 47.5 to 27.3 on the Noometry Index.

Last verified . 21 shared benchmarks.

Gemini 1.0 Pro Google

27.3

Rank #332 Confirmed

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 21 benchmarks with published results for both. Gemini 1.0 Pro scores higher in 0 categories and o3 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where o3 leads 50.2 to 9.3.
  • The biggest single-benchmark swing is MATH Level 5: 11.2% for Gemini 1.0 Pro and 97.8% for o3.

Side by side

Gemini 1.0 Pro and o3 specifications
Gemini 1.0 Proo3
ProviderGoogleOpenAI
Noometry Index27.347.5
Released2023-12-132025-04-16
WeightsProprietaryProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$2
Output $ / M tokens—$8
Results tracked2463

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3 leads

Gemini 1.0 Pro: 32.2 (#275), o3: 46.8 (#64)

Coding benchmarks
BenchmarkGemini 1.0 Proo3
LMArena Coding11081408
SWE-bench Verified—62.3%
SWE-bench Verified (bash only)—58.4%
Aider Polyglot—81.3%
GSO—8.8%
WeirdML—52.4%
CadEval—74%
ALE-Bench—933.55
HumanEval+55.5%—
MBPP+61.4%—

Agentic & Tool Use Not comparable

Gemini 1.0 Pro: —, o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkGemini 1.0 Proo3
Berkeley Function Calling Leaderboard—63%
GDPval—30.8%
DeepResearch Bench—45.2%
OSWorld—23%
LMArena Search—1144
METR Time Horizons—65.4%

Reasoning o3 leads

Gemini 1.0 Pro: 17.1 (#296), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkGemini 1.0 Proo3
LMArena Hard Prompts11091402
DTBench45.9%84.8%
Epoch Capabilities Index117.04146.86
ARC-AGI-2—6.5%
SimpleBench—53.1%
Kagi LLM Benchmark—67.6%
ARC-AGI-1—60.8%
CritPt—1.4%
Chess Puzzles—38%
EnigmaEval—13.1%
Mystery Game Puzzles—29%
LMCA—39.7%
ForecastBench—62.5

Math o3 leads

Gemini 1.0 Pro: 9.3 (#321), o3: 50.2 (#58)

Math benchmarks
BenchmarkGemini 1.0 Proo3
OTIS Mock AIME 2024-20251.1%84.4%
LMArena Math11321426
MATH Level 511.2%97.8%
FrontierMath (Tiers 1-3)—33.3%
Omni-MATH—71.4%
FrontierMath (Feb 2025 set)—18.7%
FrontierMath Tier 4 (v1)—2.1%

Knowledge o3 leads

Gemini 1.0 Pro: 15.6 (#291), o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkGemini 1.0 Proo3
GPQA Diamond34%81.8%
LMArena Expert10591402
Humanity's Last Exam—20.3%
SimpleQA Verified—49.4%
MMLU-Pro—85.9%
Confabulations—14.4%
GPQA (HELM)—75.3%
MMLU70%—

Multimodal Not comparable

Gemini 1.0 Pro: —, o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkGemini 1.0 Proo3
LMArena Vision—1214
GeoBench—74%
VPCT—52%

Multilingual o3 leads

Gemini 1.0 Pro: 33.4 (#252), o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkGemini 1.0 Proo3
LMArena Non-English11381401
LMArena Chinese11241437
LMArena French11451430
LMArena German11251420
LMArena Japanese10231403
LMArena Russian11861406
LMArena Spanish11191395
LMArena Korean—1370

Instruction Following o3 leads

Gemini 1.0 Pro: 57.6 (#267), o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkGemini 1.0 Proo3
LMArena Instruction Following11141368
IFEval—86.9%

Long Context o3 leads

Gemini 1.0 Pro: 34.3 (#249), o3: 53.3 (#6)

Long Context benchmarks
BenchmarkGemini 1.0 Proo3
LMArena Longer Query11321372
Fiction.LiveBench—88.9%
CL-bench—17.8%

Writing & Preference o3 leads

Gemini 1.0 Pro: 36.0 (#264), o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkGemini 1.0 Proo3
LMArena Text11491410
LMArena Creative Writing11311359
LMArena Multi-Turn11391405
Short-Story Creative Writing—83.9%
EQ-Bench Creative Writing—1676
WildBench—86.1%

Frequently asked questions

Is Gemini 1.0 Pro better than o3?

o3 is the stronger model overall, scoring 47.5 to 27.3 on the Noometry Index.

Is Gemini 1.0 Pro or o3 better for coding?

o3 scores higher on coding benchmarks: 46.8 versus 32.2 in the Noometry coding category.

How many benchmarks do Gemini 1.0 Pro and o3 share?

21 benchmarks have published results for both models. Gemini 1.0 Pro has 24 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper