Model comparison

GPT-4.5 vs o1-mini

GPT-4.5 is the stronger model overall, scoring 37.2 to 34.0 on the Noometry Index.

Last verified . 35 shared benchmarks.

GPT-4.5 OpenAI

37.2

Rank #208 Confirmed

o1-mini OpenAI

34.0

Rank #235 Confirmed

Summary

  • They share 35 benchmarks with published results for both. GPT-4.5 scores higher in 7 categories and o1-mini in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where GPT-4.5 leads 52.5 to 43.6.
  • The biggest single-benchmark swing is LiveBench Coding: 75.2% for GPT-4.5 and 48% for o1-mini.

Side by side

GPT-4.5 and o1-mini specifications
GPT-4.5o1-mini
ProviderOpenAIOpenAI
Noometry Index37.234.0
Released2025-02-272024-09-12
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4239

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-4.5 leads

GPT-4.5: 42.2 (#109), o1-mini: 35.5 (#224)

Coding benchmarks
BenchmarkGPT-4.5o1-mini
Aider Polyglot44.9%32.9%
WeirdML39.4%36.3%
LiveBench Coding75.2%48%
LMArena Coding13961362
HumanEval+—89%
MBPP+—78.8%

Agentic & Tool Use GPT-4.5 leads

GPT-4.5: 27.9 (#97), o1-mini: 24.6 (#118)

Agentic & Tool Use benchmarks
BenchmarkGPT-4.5o1-mini
Cybench17.5%10%

Reasoning GPT-4.5 leads

GPT-4.5: 13.9 (#330), o1-mini: 8.8 (#346)

Reasoning benchmarks
BenchmarkGPT-4.5o1-mini
ARC-AGI-20.8%0.8%
SimpleBench34.5%18.1%
ARC-AGI-110.3%14%
LiveBench Reasoning71.1%72.3%
LMArena Hard Prompts14031333
LiveBench Data Analysis64.3%57.9%
Epoch Capabilities Index136.74135.82
LiveBench69%57.8%
EnigmaEval3.2%—
ForecastBench61.7—

Math o1-mini leads

GPT-4.5: 32.6 (#211), o1-mini: 35.4 (#186)

Math benchmarks
BenchmarkGPT-4.5o1-mini
OTIS Mock AIME 2024-202537.8%46.9%
LiveBench Math69.3%62%
LMArena Math14121358
MATH Level 578.6%89.2%
FrontierMath (Feb 2025 set)—1.7%

Knowledge o1-mini leads

GPT-4.5: 32.5 (#211), o1-mini: 34.9 (#192)

Knowledge benchmarks
BenchmarkGPT-4.5o1-mini
GPQA Diamond68.7%62.4%
Confabulations13.6%18.6%
LMArena Expert13941316
Humanity's Last Exam5.4%—

Multimodal Not comparable

GPT-4.5: 37.6 (#71), o1-mini: —

Multimodal benchmarks
BenchmarkGPT-4.5o1-mini
LMArena Vision1195—
VPCT45%—

Multilingual GPT-4.5 leads

GPT-4.5: 52.5 (#83), o1-mini: 43.6 (#182)

Multilingual benchmarks
BenchmarkGPT-4.5o1-mini
LMArena Non-English14131289
LMArena Chinese14211314
LMArena French14181293
LMArena German14571278
LMArena Japanese14161245
LMArena Korean13921223
LMArena Russian14191283
LMArena Spanish—1303

Instruction Following GPT-4.5 leads

GPT-4.5: 72.6 (#134), o1-mini: 66.7 (#206)

Instruction Following benchmarks
BenchmarkGPT-4.5o1-mini
LiveBench Instruction Following72.3%65.4%
LMArena Instruction Following14041304

Long Context Too close to call

GPT-4.5: 40.4 (#155), o1-mini: 40.1 (#161)

Long Context benchmarks
BenchmarkGPT-4.5o1-mini
LMArena Longer Query14061320
Fiction.LiveBench63.9%—

Writing & Preference GPT-4.5 leads

GPT-4.5: 56.9 (#134), o1-mini: 48.4 (#202)

Writing & Preference benchmarks
BenchmarkGPT-4.5o1-mini
LMArena Text14171317
LMArena Creative Writing13941244
Short-Story Creative Writing75.6%64.9%
LMArena Multi-Turn14441314
LiveBench Language61.5%40.9%
EQ-Bench Creative Writing1258—

Frequently asked questions

Is GPT-4.5 better than o1-mini?

GPT-4.5 is the stronger model overall, scoring 37.2 to 34.0 on the Noometry Index.

Is GPT-4.5 or o1-mini better for coding?

GPT-4.5 scores higher on coding benchmarks: 42.2 versus 35.5 in the Noometry coding category.

How many benchmarks do GPT-4.5 and o1-mini share?

35 benchmarks have published results for both models. GPT-4.5 has 42 scored results on Noometry and o1-mini has 39.

Related comparisons

Go deeper