Model comparison

GPT-4.5 vs GPT-4o mini

GPT-4.5 is the stronger model overall, scoring 37.2 to 25.5 on the Noometry Index.

Last verified . 36 shared benchmarks.

GPT-4.5 OpenAI

37.2

Rank #208 Confirmed

GPT-4o mini OpenAI

25.5

Rank #343 Confirmed

Summary

  • They share 36 benchmarks with published results for both. GPT-4.5 scores higher in 10 categories and GPT-4o mini in 0 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-4.5 leads 32.6 to 10.4.
  • The biggest single-benchmark swing is Aider Polyglot: 44.9% for GPT-4.5 and 3.6% for GPT-4o mini.

Side by side

GPT-4.5 and GPT-4o mini specifications
GPT-4.5GPT-4o mini
ProviderOpenAIOpenAI
Noometry Index37.225.5
Released2025-02-272024-07-18
WeightsProprietaryProprietary
Context window—128K
Max output—16K
Input $ / M tokens—$0.15
Output $ / M tokens—$0.60
Results tracked4260

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-4.5 leads

GPT-4.5: 42.2 (#109), GPT-4o mini: 22.0 (#335)

Coding benchmarks
BenchmarkGPT-4.5GPT-4o mini
Aider Polyglot44.9%3.6%
WeirdML39.4%11.8%
LiveBench Coding75.2%43.1%
LMArena Coding13961290
BigCodeBench Instruct—46.1%
BigCodeBench Complete—57.4%
HumanEval+—83.5%
MBPP+—72.2%

Agentic & Tool Use Too close to call

GPT-4.5: 27.9 (#97), GPT-4o mini: 27.5 (#101)

Agentic & Tool Use benchmarks
BenchmarkGPT-4.5GPT-4o mini
Cybench17.5%—
BALROG—17.4%

Reasoning GPT-4.5 leads

GPT-4.5: 13.9 (#330), GPT-4o mini: 8.7 (#347)

Reasoning benchmarks
BenchmarkGPT-4.5GPT-4o mini
ARC-AGI-20.8%0%
SimpleBench34.5%10.7%
LiveBench Reasoning71.1%32.8%
LMArena Hard Prompts14031267
LiveBench Data Analysis64.3%50%
Epoch Capabilities Index136.74126.56
LiveBench69%41.3%
Kagi LLM Benchmark—28.8%
ARC-AGI-110.3%—
Chess Puzzles—0%
EnigmaEval3.2%—
Mystery Game Puzzles—12%
DTBench—54.4%
LMCA—10.4%
ForecastBench61.7—
PIQA—88.7%

Math GPT-4.5 leads

GPT-4.5: 32.6 (#211), GPT-4o mini: 10.4 (#314)

Math benchmarks
BenchmarkGPT-4.5GPT-4o mini
OTIS Mock AIME 2024-202537.8%6.9%
LiveBench Math69.3%36.3%
LMArena Math14121267
MATH Level 578.6%52.6%
FrontierMath (Tiers 1-3)—0.7%
Omni-MATH—28%
GSM8K—91.3%

Knowledge GPT-4.5 leads

GPT-4.5: 32.5 (#211), GPT-4o mini: 17.7 (#284)

Knowledge benchmarks
BenchmarkGPT-4.5GPT-4o mini
GPQA Diamond68.7%37.7%
Confabulations13.6%37.2%
LMArena Expert13941235
Humanity's Last Exam5.4%—
SimpleQA Verified—8.3%
MMLU-Pro—60.3%
GPQA (HELM)—36.8%
BoolQ—88.7%
MMLU—81.8%

Multimodal GPT-4.5 leads

GPT-4.5: 37.6 (#71), GPT-4o mini: 25.9 (#122)

Multimodal benchmarks
BenchmarkGPT-4.5GPT-4o mini
LMArena Vision11951066
VPCT45%34%
Video-MME—64.8%
GeoBench—64%

Multilingual GPT-4.5 leads

GPT-4.5: 52.5 (#83), GPT-4o mini: 42.0 (#199)

Multilingual benchmarks
BenchmarkGPT-4.5GPT-4o mini
LMArena Non-English14131266
LMArena Chinese14211265
LMArena French14181297
LMArena German14571272
LMArena Japanese14161216
LMArena Korean13921195
LMArena Russian14191275
LMArena Spanish—1276

Instruction Following GPT-4.5 leads

GPT-4.5: 72.6 (#134), GPT-4o mini: 61.9 (#239)

Instruction Following benchmarks
BenchmarkGPT-4.5GPT-4o mini
LiveBench Instruction Following72.3%56.8%
LMArena Instruction Following14041258
IFEval—78.2%

Long Context GPT-4.5 leads

GPT-4.5: 40.4 (#155), GPT-4o mini: 39.1 (#186)

Long Context benchmarks
BenchmarkGPT-4.5GPT-4o mini
LMArena Longer Query14061289
Fiction.LiveBench63.9%—

Writing & Preference GPT-4.5 leads

GPT-4.5: 56.9 (#134), GPT-4o mini: 39.5 (#248)

Writing & Preference benchmarks
BenchmarkGPT-4.5GPT-4o mini
LMArena Text14171286
LMArena Creative Writing13941268
Short-Story Creative Writing75.6%67.2%
EQ-Bench Creative Writing1258873
LMArena Multi-Turn14441285
LiveBench Language61.5%28.6%
WildBench—79.1%

Frequently asked questions

Is GPT-4.5 better than GPT-4o mini?

GPT-4.5 is the stronger model overall, scoring 37.2 to 25.5 on the Noometry Index.

Is GPT-4.5 or GPT-4o mini better for coding?

GPT-4.5 scores higher on coding benchmarks: 42.2 versus 22.0 in the Noometry coding category.

How many benchmarks do GPT-4.5 and GPT-4o mini share?

36 benchmarks have published results for both models. GPT-4.5 has 42 scored results on Noometry and GPT-4o mini has 60.

Related comparisons

Go deeper