Model comparison

GPT-4.5 vs o3-pro

o3-pro is the stronger model overall, scoring 42.9 to 37.2 on the Noometry Index.

Last verified . 8 shared benchmarks.

GPT-4.5 OpenAI

37.2

Rank #208 Confirmed

o3-pro OpenAI

42.9

Rank #105 Confirmed

Summary

  • They share 8 benchmarks with published results for both. GPT-4.5 scores higher in 1 category and o3-pro in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in long context, where o3-pro leads 72.2 to 40.4.
  • The biggest single-benchmark swing is ARC-AGI-1: 10.3% for GPT-4.5 and 59.3% for o3-pro.

Side by side

GPT-4.5 and o3-pro specifications
GPT-4.5o3-pro
ProviderOpenAIOpenAI
Noometry Index37.242.9
Released2025-02-272025-06-10
WeightsProprietaryProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$20
Output $ / M tokens—$80
Results tracked4212

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3-pro leads

GPT-4.5: 42.2 (#109), o3-pro: 55.5 (#24)

Coding benchmarks
BenchmarkGPT-4.5o3-pro
Aider Polyglot44.9%84.9%
WeirdML39.4%58.2%
LiveBench Coding75.2%—
LMArena Coding1396—

Agentic & Tool Use Not comparable

GPT-4.5: 27.9 (#97), o3-pro: —

Agentic & Tool Use benchmarks
BenchmarkGPT-4.5o3-pro
Cybench17.5%—

Reasoning o3-pro leads

GPT-4.5: 13.9 (#330), o3-pro: 23.8 (#171)

Reasoning benchmarks
BenchmarkGPT-4.5o3-pro
ARC-AGI-20.8%4.9%
ARC-AGI-110.3%59.3%
Epoch Capabilities Index136.74147.42
SimpleBench34.5%—
Kagi LLM Benchmark—72.1%
EnigmaEval3.2%—
LiveBench Reasoning71.1%—
LMArena Hard Prompts1403—
DTBench—86.9%
LiveBench Data Analysis64.3%—
LMCA—38.5%
ForecastBench61.7—
LiveBench69%—

Math Not comparable

GPT-4.5: 32.6 (#211), o3-pro: —

Math benchmarks
BenchmarkGPT-4.5o3-pro
OTIS Mock AIME 2024-202537.8%—
LiveBench Math69.3%—
LMArena Math1412—
MATH Level 578.6%—

Knowledge GPT-4.5 leads

GPT-4.5: 32.5 (#211), o3-pro: 29.5 (#238)

Knowledge benchmarks
BenchmarkGPT-4.5o3-pro
Confabulations13.6%14.2%
GPQA Diamond68.7%—
Humanity's Last Exam5.4%—
Vectara Hallucination Rate—23.3%
LMArena Expert1394—

Multimodal Not comparable

GPT-4.5: 37.6 (#71), o3-pro: —

Multimodal benchmarks
BenchmarkGPT-4.5o3-pro
LMArena Vision1195—
VPCT45%—

Multilingual Not comparable

GPT-4.5: 52.5 (#83), o3-pro: —

Multilingual benchmarks
BenchmarkGPT-4.5o3-pro
LMArena Non-English1413—
LMArena Chinese1421—
LMArena French1418—
LMArena German1457—
LMArena Japanese1416—
LMArena Korean1392—
LMArena Russian1419—

Instruction Following Not comparable

GPT-4.5: 72.6 (#134), o3-pro: —

Instruction Following benchmarks
BenchmarkGPT-4.5o3-pro
LiveBench Instruction Following72.3%—
LMArena Instruction Following1404—

Long Context o3-pro leads

GPT-4.5: 40.4 (#155), o3-pro: 72.2 (#1)

Long Context benchmarks
BenchmarkGPT-4.5o3-pro
Fiction.LiveBench63.9%97.2%
LMArena Longer Query1406—

Writing & Preference Too close to call

GPT-4.5: 56.9 (#134), o3-pro: 57.1 (#133)

Writing & Preference benchmarks
BenchmarkGPT-4.5o3-pro
Short-Story Creative Writing75.6%84.4%
LMArena Text1417—
LMArena Creative Writing1394—
EQ-Bench Creative Writing1258—
LMArena Multi-Turn1444—
LiveBench Language61.5%—

Frequently asked questions

Is GPT-4.5 better than o3-pro?

o3-pro is the stronger model overall, scoring 42.9 to 37.2 on the Noometry Index.

Is GPT-4.5 or o3-pro better for coding?

o3-pro scores higher on coding benchmarks: 55.5 versus 42.2 in the Noometry coding category.

How many benchmarks do GPT-4.5 and o3-pro share?

8 benchmarks have published results for both models. GPT-4.5 has 42 scored results on Noometry and o3-pro has 12.

Related comparisons

Go deeper