Model comparison

GPT-4o vs GPT-5.4

GPT-5.4 is the stronger model overall, scoring 59.4 to 28.6 on the Noometry Index.

Last verified . 39 shared benchmarks.

GPT-4o OpenAI

28.6

Rank #324 Confirmed

GPT-5.4 OpenAI

59.4

Rank #16 Confirmed

Summary

  • They share 39 benchmarks with published results for both. GPT-4o scores higher in 0 categories and GPT-5.4 in 10 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-5.4 leads 73.5 to 10.6.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 6.4% for GPT-4o and 97.8% for GPT-5.4.
  • GPT-4o is cheaper at $2.50 / $10 per million input/output tokens, against $2.50 / $15 for GPT-5.4.
  • GPT-5.4 accepts more context: 1.05M tokens versus 128K.

Side by side

GPT-4o and GPT-5.4 specifications
GPT-4oGPT-5.4
ProviderOpenAIOpenAI
Noometry Index28.659.4
Released2024-05-132026-03-05
WeightsProprietaryProprietary
Context window128K1.05M
Max output16K128K
Input $ / M tokens$2.50$2.50
Output $ / M tokens$10$15
Results tracked7268

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.4 leads

GPT-4o: 24.8 (#328), GPT-5.4: 52.6 (#33)

Coding benchmarks
BenchmarkGPT-4oGPT-5.4
SWE-bench Verified31%76.9%
GSO0%31.4%
WeirdML25.1%77.7%
LMArena Coding12971497
DeepSWE—51.8%
SWE-bench Verified (bash only)21.6%—
Aider Polyglot45.3%—
LMArena WebDev—1465
SciCode—56.6%
BigCodeBench Instruct51.1%—
LiveBench Coding51.4%—
MirrorCode—15.6%
BigCodeBench Complete61.1%—
CadEval26%—
ALE-Bench—1,607
AlgoTune—1.85
HumanEval+87.2%—
MBPP+72.2%—

Agentic & Tool Use GPT-5.4 leads

GPT-4o: 21.0 (#141), GPT-5.4: 46.5 (#13)

Agentic & Tool Use benchmarks
BenchmarkGPT-4oGPT-5.4
LMArena Search10061197
METR Time Horizons40.8%74.3%
Terminal-Bench—81.8%
APEX-Agents—52.4%
GDPval9.9%—
TheAgentCompany8.6%—
τ²-bench Banking—39.4%
Cybench12.5%—
DeepResearch Bench—35.1%
PostTrainBench—19%
BALROG32.3%—
GBAEval—45.1%
Vending-Bench 2—6,144

Reasoning GPT-5.4 leads

GPT-4o: 9.4 (#343), GPT-5.4: 61.8 (#19)

Reasoning benchmarks
BenchmarkGPT-4oGPT-5.4
ARC-AGI-20%74%
ARC-AGI-14.5%93.7%
CritPt0%23.4%
Chess Puzzles13%44%
EnigmaEval0.8%16%
LMArena Hard Prompts12811485
DTBench64.5%94.4%
LMCA16.6%52%
Epoch Capabilities Index128.97156.81
ForecastBench57.759.5
SimpleBench17.8%—
Kagi LLM Benchmark—63.8%
NYT Connections (extended)—91.3%
Thematic Generalization—80%
EBR-Bench—25.4%
LiveBench Reasoning55.8%—
Mystery Game Puzzles—37%
LiveBench Data Analysis60.9%—
LiveBench55.3%—

Math GPT-5.4 leads

GPT-4o: 10.6 (#312), GPT-5.4: 73.5 (#19)

Knowledge GPT-5.4 leads

GPT-4o: 28.8 (#242), GPT-5.4: 65.3 (#14)

Knowledge benchmarks
BenchmarkGPT-4oGPT-5.4
GPQA Diamond49.2%93.3%
Humanity's Last Exam2.7%36.2%
SimpleQA Verified26%45.1%
Vectara Hallucination Rate9.6%7%
LMArena Expert12501507
MMLU-Pro71.3%—
Confabulations15.3%—
GPQA (HELM)52%—
MMLU88.1%—

Multimodal GPT-5.4 leads

GPT-4o: 34.5 (#91), GPT-5.4: 43.7 (#20)

Multimodal benchmarks
BenchmarkGPT-4oGPT-5.4
LMArena Vision11371303
Video-MME71.9%—
GeoBench71%—
VPCT40%—
Blueprint-Bench 2—27.1%
Furniture Assembly—37.5%
LMArena Document—1471
ScienceQA88.5%—

Multilingual GPT-5.4 leads

GPT-4o: 43.2 (#186), GPT-5.4: 56.2 (#23)

Multilingual benchmarks
BenchmarkGPT-4oGPT-5.4
LMArena Non-English12831465
LMArena Chinese12771519
LMArena French13041493
LMArena German12821472
LMArena Japanese12571485
LMArena Korean12341448
LMArena Russian12861480
LMArena Spanish12921454

Instruction Following GPT-5.4 leads

GPT-4o: 66.6 (#207), GPT-5.4: 77.1 (#27)

Instruction Following benchmarks
BenchmarkGPT-4oGPT-5.4
LMArena Instruction Following12781469
LiveBench Instruction Following68.6%—
IFEval81.7%—

Long Context GPT-5.4 leads

GPT-4o: 39.4 (#179), GPT-5.4: 50.3 (#8)

Long Context benchmarks
BenchmarkGPT-4oGPT-5.4
LMArena Longer Query12891473
Fiction.LiveBench66.7%—
CL-bench—27.9%
CL-bench Life—21.7%

Writing & Preference GPT-5.4 leads

GPT-4o: 52.6 (#166), GPT-5.4: 71.9 (#17)

Writing & Preference benchmarks
BenchmarkGPT-4oGPT-5.4
LMArena Text13001469
LMArena Creative Writing12921439
LMArena Multi-Turn13021482
Short-Story Creative Writing81.8%—
EQ-Bench Creative Writing—1840
WildBench82.8%—
EQ-Bench 4—1272
LiveBench Language47.6%—

Frequently asked questions

Is GPT-4o better than GPT-5.4?

GPT-5.4 is the stronger model overall, scoring 59.4 to 28.6 on the Noometry Index.

Which is cheaper, GPT-4o or GPT-5.4?

GPT-4o is cheaper. It lists at $2.50 per million input tokens and $10 per million output tokens; GPT-5.4 lists at $2.50 and $15.

Is GPT-4o or GPT-5.4 better for coding?

GPT-5.4 scores higher on coding benchmarks: 52.6 versus 24.8 in the Noometry coding category.

Which has the bigger context window?

GPT-5.4 does, with 1.05M tokens against 128K.

How many benchmarks do GPT-4o and GPT-5.4 share?

39 benchmarks have published results for both models. GPT-4o has 72 scored results on Noometry and GPT-5.4 has 68.

Related comparisons

Go deeper