Model comparison

GPT-4.5 vs GPT-5.4 mini

GPT-5.4 mini is the stronger model overall, scoring 45.0 to 37.2 on the Noometry Index.

Last verified . 25 shared benchmarks.

GPT-4.5 OpenAI

37.2

Rank #208 Confirmed

GPT-5.4 mini OpenAI

45.0

Rank #76 Confirmed

Summary

  • They share 25 benchmarks with published results for both. GPT-4.5 scores higher in 1 category and GPT-5.4 mini in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where GPT-5.4 mini leads 51.5 to 32.5.
  • The biggest single-benchmark swing is ARC-AGI-1: 10.3% for GPT-4.5 and 63.7% for GPT-5.4 mini.

Side by side

GPT-4.5 and GPT-5.4 mini specifications
GPT-4.5GPT-5.4 mini
ProviderOpenAIOpenAI
Noometry Index37.245.0
Released2025-02-272026-03-17
WeightsProprietaryProprietary
Context window—400K
Max output—128K
Input $ / M tokens—$0.75
Output $ / M tokens—$4.50
Results tracked4246

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.4 mini leads

GPT-4.5: 42.2 (#109), GPT-5.4 mini: 45.2 (#72)

Coding benchmarks
BenchmarkGPT-4.5GPT-5.4 mini
WeirdML39.4%60.3%
LMArena Coding13961438
FrontierCode—27%
Aider Polyglot44.9%—
LMArena WebDev—1397
SciCode—49.9%
LiveBench Coding75.2%—
ALE-Bench—1,189

Agentic & Tool Use GPT-5.4 mini leads

GPT-4.5: 27.9 (#97), GPT-5.4 mini: 29.9 (#81)

Agentic & Tool Use benchmarks
BenchmarkGPT-4.5GPT-5.4 mini
Cybench17.5%—
DeepResearch Bench—36.3%

Reasoning GPT-5.4 mini leads

GPT-4.5: 13.9 (#330), GPT-5.4 mini: 30.4 (#85)

Reasoning benchmarks
BenchmarkGPT-4.5GPT-5.4 mini
ARC-AGI-20.8%18.9%
ARC-AGI-110.3%63.7%
LMArena Hard Prompts14031424
Epoch Capabilities Index136.74148.84
ForecastBench61.757
SimpleBench34.5%—
Kagi LLM Benchmark—37.9%
NYT Connections (extended)—61.8%
CritPt—10%
Chess Puzzles—24%
EnigmaEval3.2%—
Thematic Generalization—61.7%
LiveBench Reasoning71.1%—
Mystery Game Puzzles—11%
DTBench—80%
LiveBench Data Analysis64.3%—
LMCA—40.8%
LiveBench69%—

Math GPT-5.4 mini leads

GPT-4.5: 32.6 (#211), GPT-5.4 mini: 45.5 (#75)

Math benchmarks
BenchmarkGPT-4.5GPT-5.4 mini
OTIS Mock AIME 2024-202537.8%88.9%
LMArena Math14121419
FrontierMath (Tiers 1-3)—51.2%
FrontierMath Tier 4—9.8%
ProofBench—21%
LiveBench Math69.3%—
MATH Level 578.6%—
FrontierMath (Feb 2025 set)—28.3%
FrontierMath Tier 4 (v1)—2.1%

Knowledge GPT-5.4 mini leads

GPT-4.5: 32.5 (#211), GPT-5.4 mini: 51.5 (#67)

Knowledge benchmarks
BenchmarkGPT-4.5GPT-5.4 mini
GPQA Diamond68.7%86.9%
LMArena Expert13941435
Humanity's Last Exam5.4%—
SimpleQA Verified—29.4%
Confabulations13.6%—
Vectara Hallucination Rate—5.5%

Multimodal GPT-5.4 mini leads

GPT-4.5: 37.6 (#71), GPT-5.4 mini: 39.7 (#56)

Multimodal benchmarks
BenchmarkGPT-4.5GPT-5.4 mini
LMArena Vision11951245
VPCT45%—

Multilingual Too close to call

GPT-4.5: 52.5 (#83), GPT-5.4 mini: 51.9 (#96)

Multilingual benchmarks
BenchmarkGPT-4.5GPT-5.4 mini
LMArena Non-English14131405
LMArena Chinese14211446
LMArena French14181440
LMArena German14571409
LMArena Japanese14161374
LMArena Korean13921368
LMArena Russian14191417
LMArena Spanish—1405

Instruction Following GPT-5.4 mini leads

GPT-4.5: 72.6 (#134), GPT-5.4 mini: 74.1 (#102)

Instruction Following benchmarks
BenchmarkGPT-4.5GPT-5.4 mini
LMArena Instruction Following14041405
LiveBench Instruction Following72.3%—

Long Context GPT-5.4 mini leads

GPT-4.5: 40.4 (#155), GPT-5.4 mini: 43.0 (#112)

Long Context benchmarks
BenchmarkGPT-4.5GPT-5.4 mini
LMArena Longer Query14061407
Fiction.LiveBench63.9%—

Writing & Preference GPT-5.4 mini leads

GPT-4.5: 56.9 (#134), GPT-5.4 mini: 64.0 (#58)

Writing & Preference benchmarks
BenchmarkGPT-4.5GPT-5.4 mini
LMArena Text14171412
LMArena Creative Writing13941370
EQ-Bench Creative Writing12581665
LMArena Multi-Turn14441429
Short-Story Creative Writing75.6%—
LiveBench Language61.5%—

Frequently asked questions

Is GPT-4.5 better than GPT-5.4 mini?

GPT-5.4 mini is the stronger model overall, scoring 45.0 to 37.2 on the Noometry Index.

Is GPT-4.5 or GPT-5.4 mini better for coding?

GPT-5.4 mini scores higher on coding benchmarks: 45.2 versus 42.2 in the Noometry coding category.

How many benchmarks do GPT-4.5 and GPT-5.4 mini share?

25 benchmarks have published results for both models. GPT-4.5 has 42 scored results on Noometry and GPT-5.4 mini has 46.

Related comparisons

Go deeper