Model comparison

Gemini 1.5 Pro (May 2024) vs o1-pro

Gemini 1.5 Pro (May 2024) and o1-pro score almost the same on the Noometry Index (32.1 vs 31.5), so choose on price, context window or the category you care about most.

Last verified . 1 shared benchmarks.

Gemini 1.5 Pro (May 2024) Google

32.1

Rank #261 Confirmed

o1-pro OpenAI

31.5

Rank #271 Reported

Summary

  • They share 1 benchmark with published results for both. Gemini 1.5 Pro (May 2024) scores higher in 0 categories and o1-pro in 2 categories; one gap is clear of the uncertainty.
  • The widest gap is in reasoning, where o1-pro leads 20.4 to 12.3.

Side by side

Gemini 1.5 Pro (May 2024) and o1-pro specifications
Gemini 1.5 Pro (May 2024)o1-pro
ProviderGoogleOpenAI
Noometry Index32.131.5
Released2024-02-152025-03-19
WeightsProprietaryProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$150
Output $ / M tokens—$600
Results tracked453

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Gemini 1.5 Pro (May 2024): 34.2 (#241), o1-pro: —

Coding benchmarks
BenchmarkGemini 1.5 Pro (May 2024)o1-pro
WeirdML22.2%—
BigCodeBench Instruct43.8%—
LMArena Coding1294—
BigCodeBench Complete57.5%—
CadEval34%—
HumanEval+79.3%—
MBPP+74.6%—

Agentic & Tool Use Not comparable

Gemini 1.5 Pro (May 2024): 17.9 (#145), o1-pro: —

Agentic & Tool Use benchmarks
BenchmarkGemini 1.5 Pro (May 2024)o1-pro
TheAgentCompany3.4%—
Cybench7.5%—
BALROG21%—

Reasoning o1-pro leads

Gemini 1.5 Pro (May 2024): 12.3 (#338), o1-pro: 20.4 (#239)

Reasoning benchmarks
BenchmarkGemini 1.5 Pro (May 2024)o1-pro
ARC-AGI-20.8%—
SimpleBench27.1%—
ARC-AGI-1—23.3%
EnigmaEval—6.1%
LMArena Hard Prompts1296—
DTBench59%—
BIG-Bench Hard89.2%—
Epoch Capabilities Index131.73—
ForecastBench58.4—

Math Not comparable

Gemini 1.5 Pro (May 2024): 25.8 (#266), o1-pro: —

Math benchmarks
BenchmarkGemini 1.5 Pro (May 2024)o1-pro
OTIS Mock AIME 2024-202523.1%—
Omni-MATH36.4%—
LMArena Math1315—
MATH Level 570.4%—

Knowledge Too close to call

Gemini 1.5 Pro (May 2024): 29.4 (#239), o1-pro: 29.7 (#234)

Knowledge benchmarks
BenchmarkGemini 1.5 Pro (May 2024)o1-pro
Humanity's Last Exam4.6%8.1%
GPQA Diamond57.2%—
MMLU-Pro73.7%—
Confabulations13.5%—
GPQA (HELM)53.4%—
LMArena Expert1279—
MMLU86.9%—

Multimodal Not comparable

Gemini 1.5 Pro (May 2024): 36.8 (#77), o1-pro: —

Multimodal benchmarks
BenchmarkGemini 1.5 Pro (May 2024)o1-pro
LMArena Vision1161—
Video-MME75%—

Multilingual Not comparable

Gemini 1.5 Pro (May 2024): 45.3 (#174), o1-pro: —

Multilingual benchmarks
BenchmarkGemini 1.5 Pro (May 2024)o1-pro
LMArena Non-English1312—
LMArena Chinese1331—
LMArena French1302—
LMArena German1286—
LMArena Japanese1292—
LMArena Korean1298—
LMArena Russian1320—
LMArena Spanish1311—

Instruction Following Not comparable

Gemini 1.5 Pro (May 2024): 68.6 (#185), o1-pro: —

Instruction Following benchmarks
BenchmarkGemini 1.5 Pro (May 2024)o1-pro
IFEval83.7%—
LMArena Instruction Following1297—

Long Context Not comparable

Gemini 1.5 Pro (May 2024): 39.8 (#169), o1-pro: —

Long Context benchmarks
BenchmarkGemini 1.5 Pro (May 2024)o1-pro
LMArena Longer Query1308—

Writing & Preference Not comparable

Gemini 1.5 Pro (May 2024): 52.4 (#172), o1-pro: —

Writing & Preference benchmarks
BenchmarkGemini 1.5 Pro (May 2024)o1-pro
LMArena Text1319—
LMArena Creative Writing1333—
WildBench81.3%—
LMArena Multi-Turn1296—

Frequently asked questions

Is Gemini 1.5 Pro (May 2024) better than o1-pro?

Gemini 1.5 Pro (May 2024) and o1-pro score almost the same on the Noometry Index (32.1 vs 31.5), so choose on price, context window or the category you care about most.

How many benchmarks do Gemini 1.5 Pro (May 2024) and o1-pro share?

1 benchmark has published results for both models. Gemini 1.5 Pro (May 2024) has 45 scored results on Noometry and o1-pro has 3.

Related comparisons

Go deeper