Model comparison

GPT-4.5 vs o3

o3 is the stronger model overall, scoring 47.5 to 37.2 on the Noometry Index.

Last verified . 34 shared benchmarks.

GPT-4.5 OpenAI

37.2

Rank #208 Confirmed

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 34 benchmarks with published results for both. GPT-4.5 scores higher in 1 category and o3 in 9 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where o3 leads 54.6 to 32.5.
  • The biggest single-benchmark swing is ARC-AGI-1: 10.3% for GPT-4.5 and 60.8% for o3.

Side by side

GPT-4.5 and o3 specifications
GPT-4.5o3
ProviderOpenAIOpenAI
Noometry Index37.247.5
Released2025-02-272025-04-16
WeightsProprietaryProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$2
Output $ / M tokens—$8
Results tracked4263

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3 leads

GPT-4.5: 42.2 (#109), o3: 46.8 (#64)

Coding benchmarks
BenchmarkGPT-4.5o3
Aider Polyglot44.9%81.3%
WeirdML39.4%52.4%
LMArena Coding13961408
SWE-bench Verified—62.3%
SWE-bench Verified (bash only)—58.4%
GSO—8.8%
LiveBench Coding75.2%—
CadEval—74%
ALE-Bench—933.55

Agentic & Tool Use o3 leads

GPT-4.5: 27.9 (#97), o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkGPT-4.5o3
Berkeley Function Calling Leaderboard—63%
GDPval—30.8%
Cybench17.5%—
DeepResearch Bench—45.2%
OSWorld—23%
LMArena Search—1144
METR Time Horizons—65.4%

Reasoning o3 leads

GPT-4.5: 13.9 (#330), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkGPT-4.5o3
ARC-AGI-20.8%6.5%
SimpleBench34.5%53.1%
ARC-AGI-110.3%60.8%
EnigmaEval3.2%13.1%
LMArena Hard Prompts14031402
Epoch Capabilities Index136.74146.86
ForecastBench61.762.5
Kagi LLM Benchmark—67.6%
CritPt—1.4%
Chess Puzzles—38%
LiveBench Reasoning71.1%—
Mystery Game Puzzles—29%
DTBench—84.8%
LiveBench Data Analysis64.3%—
LMCA—39.7%
LiveBench69%—

Math o3 leads

GPT-4.5: 32.6 (#211), o3: 50.2 (#58)

Math benchmarks
BenchmarkGPT-4.5o3
OTIS Mock AIME 2024-202537.8%84.4%
LMArena Math14121426
MATH Level 578.6%97.8%
FrontierMath (Tiers 1-3)—33.3%
Omni-MATH—71.4%
LiveBench Math69.3%—
FrontierMath (Feb 2025 set)—18.7%
FrontierMath Tier 4 (v1)—2.1%

Knowledge o3 leads

GPT-4.5: 32.5 (#211), o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkGPT-4.5o3
GPQA Diamond68.7%81.8%
Humanity's Last Exam5.4%20.3%
Confabulations13.6%14.4%
LMArena Expert13941402
SimpleQA Verified—49.4%
MMLU-Pro—85.9%
GPQA (HELM)—75.3%

Multimodal o3 leads

GPT-4.5: 37.6 (#71), o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkGPT-4.5o3
LMArena Vision11951214
VPCT45%52%
GeoBench—74%

Multilingual Too close to call

GPT-4.5: 52.5 (#83), o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkGPT-4.5o3
LMArena Non-English14131401
LMArena Chinese14211437
LMArena French14181430
LMArena German14571420
LMArena Japanese14161403
LMArena Korean13921370
LMArena Russian14191406
LMArena Spanish—1395

Instruction Following Too close to call

GPT-4.5: 72.6 (#134), o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkGPT-4.5o3
LMArena Instruction Following14041368
LiveBench Instruction Following72.3%—
IFEval—86.9%

Long Context o3 leads

GPT-4.5: 40.4 (#155), o3: 53.3 (#6)

Long Context benchmarks
BenchmarkGPT-4.5o3
Fiction.LiveBench63.9%88.9%
LMArena Longer Query14061372
CL-bench—17.8%

Writing & Preference o3 leads

GPT-4.5: 56.9 (#134), o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkGPT-4.5o3
LMArena Text14171410
LMArena Creative Writing13941359
Short-Story Creative Writing75.6%83.9%
EQ-Bench Creative Writing12581676
LMArena Multi-Turn14441405
WildBench—86.1%
LiveBench Language61.5%—

Frequently asked questions

Is GPT-4.5 better than o3?

o3 is the stronger model overall, scoring 47.5 to 37.2 on the Noometry Index.

Is GPT-4.5 or o3 better for coding?

o3 scores higher on coding benchmarks: 46.8 versus 42.2 in the Noometry coding category.

How many benchmarks do GPT-4.5 and o3 share?

34 benchmarks have published results for both models. GPT-4.5 has 42 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper