Model comparison

o1-mini vs o1-pro

o1-mini is the stronger model overall, scoring 34.0 to 31.5 on the Noometry Index.

Last verified . 1 shared benchmarks.

o1-mini OpenAI

34.0

Rank #235 Confirmed

o1-pro OpenAI

31.5

Rank #271 Reported

Summary

  • They share 1 benchmark with published results for both. o1-mini scores higher in 1 category and o1-pro in 1 category; 2 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where o1-pro leads 20.4 to 8.8.
  • The biggest single-benchmark swing is ARC-AGI-1: 14% for o1-mini and 23.3% for o1-pro.

Side by side

o1-mini and o1-pro specifications
o1-minio1-pro
ProviderOpenAIOpenAI
Noometry Index34.031.5
Released2024-09-122025-03-19
WeightsProprietaryProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$150
Output $ / M tokens—$600
Results tracked393

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

o1-mini: 35.5 (#224), o1-pro: —

Coding benchmarks
Benchmarko1-minio1-pro
Aider Polyglot32.9%—
WeirdML36.3%—
LiveBench Coding48%—
LMArena Coding1362—
HumanEval+89%—
MBPP+78.8%—

Agentic & Tool Use Not comparable

o1-mini: 24.6 (#118), o1-pro: —

Agentic & Tool Use benchmarks
Benchmarko1-minio1-pro
Cybench10%—

Reasoning o1-pro leads

o1-mini: 8.8 (#346), o1-pro: 20.4 (#239)

Reasoning benchmarks
Benchmarko1-minio1-pro
ARC-AGI-114%23.3%
ARC-AGI-20.8%—
SimpleBench18.1%—
EnigmaEval—6.1%
LiveBench Reasoning72.3%—
LMArena Hard Prompts1333—
LiveBench Data Analysis57.9%—
Epoch Capabilities Index135.82—
LiveBench57.8%—

Math Not comparable

o1-mini: 35.4 (#186), o1-pro: —

Math benchmarks
Benchmarko1-minio1-pro
OTIS Mock AIME 2024-202546.9%—
LiveBench Math62%—
LMArena Math1358—
MATH Level 589.2%—
FrontierMath (Feb 2025 set)1.7%—

Knowledge o1-mini leads

o1-mini: 34.9 (#192), o1-pro: 29.7 (#234)

Knowledge benchmarks
Benchmarko1-minio1-pro
GPQA Diamond62.4%—
Humanity's Last Exam—8.1%
Confabulations18.6%—
LMArena Expert1316—

Multilingual Not comparable

o1-mini: 43.6 (#182), o1-pro: —

Multilingual benchmarks
Benchmarko1-minio1-pro
LMArena Non-English1289—
LMArena Chinese1314—
LMArena French1293—
LMArena German1278—
LMArena Japanese1245—
LMArena Korean1223—
LMArena Russian1283—
LMArena Spanish1303—

Instruction Following Not comparable

o1-mini: 66.7 (#206), o1-pro: —

Instruction Following benchmarks
Benchmarko1-minio1-pro
LiveBench Instruction Following65.4%—
LMArena Instruction Following1304—

Long Context Not comparable

o1-mini: 40.1 (#161), o1-pro: —

Long Context benchmarks
Benchmarko1-minio1-pro
LMArena Longer Query1320—

Writing & Preference Not comparable

o1-mini: 48.4 (#202), o1-pro: —

Writing & Preference benchmarks
Benchmarko1-minio1-pro
LMArena Text1317—
LMArena Creative Writing1244—
Short-Story Creative Writing64.9%—
LMArena Multi-Turn1314—
LiveBench Language40.9%—

Frequently asked questions

Is o1-mini better than o1-pro?

o1-mini is the stronger model overall, scoring 34.0 to 31.5 on the Noometry Index.

How many benchmarks do o1-mini and o1-pro share?

1 benchmark has published results for both models. o1-mini has 39 scored results on Noometry and o1-pro has 3.

Related comparisons

Go deeper