OpenAI, proprietary

GPT-5.4

GPT-5.4 by OpenAI ranks 16th of 354 ranked models on the Noometry Index as of October 2026, with a score of 59.4. Its strongest category is long context, where it ranks 8th. API pricing starts at $2.50 per million input tokens and $15 per million output tokens, with a 1.05M-token context window.

Last verified

Specifications

Noometry rank
#16 of 354
Index score
59.4
Evidence
Confirmed 68 results
Provider
OpenAI
Released
March 5, 2026
Weights
Proprietary
Reasoning
Yes
Context window
1.05M
Max output
128K
Input price
$2.50 / M
Output price
$15 / M
Blended price
$5.63 / M
Output speed
12 tokens/s Kagi
Value
#183 of 219
Knowledge cutoff
August 2025
Input
text, image, pdf

Category scores

Each category score combines every public result we have in that category.

GPT-5.4 category scores
  1. Coding 52.6
  2. Agentic & Tool Use 46.5
  3. Reasoning 61.8
  4. Math 73.5
  5. Knowledge 65.3
  6. Multimodal 43.7
  7. Multilingual 56.2
  8. Instruction Following 77.1
  9. Long Context 50.3
  10. Writing & Preference 71.9
GPT-5.4 category ranks
CategoryScoreRankResults
Coding52.6#338
Agentic & Tool Use46.5#136
Reasoning61.8#1913
Math73.5#196
Knowledge65.3#145
Multimodal43.7#203
Multilingual56.2#231
Instruction Following77.1#271
Long Context50.3#83
Writing & Preference71.9#175

Strengths and weaknesses

Categories where GPT-5.4 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

GPT-5.4: strongest categories
CategoryScorevs medianRank
Long Context50.3+9.4#8 of 296, top 3%
Knowledge65.3+28.0#14 of 314, top 5%
Reasoning61.8+38.2#19 of 350, top 6%

Weakest categories

GPT-5.4: weakest categories
CategoryScorevs medianRank
Multimodal43.7+5.1#20 of 128, top 16%
Coding52.6+13.9#33 of 340, top 10%
Instruction Following77.1+5.9#27 of 305, top 9%

Closest competitors

The models ranked just above and below GPT-5.4. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to GPT-5.4
ModelRankScoreBlended $/MSpeed
GPT-6 Sol#1261.8$4—Compare
Claude Opus 4.8#1360.7$1034Compare
Gemini 3.7 Flash#1459.8$1.50—Compare
Kimi K3#1559.5$6—Compare
GPT-5.6 Terra#1759.2$4.5011Compare
GPT-5.4 Pro#1858.9$67.50—Compare
Claude Opus 4.7#1958.3$1033Compare
Claude Opus 4.6#2058.2$1019Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

GPT-5.4 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified76.9%#8 of 32, top 25%highEpoch AI2026-03-06
DeepSWE51.8%#23 of 29, top 80%xhighEpoch AI
LMArena WebDev1465#54 of 113, top 48%LMArena2026-10-08
LMArena WebDev1400LMArena2026-10-08
LMArena WebDev1444LMArena2026-10-08
SciCode56.6%#18 of 121, top 15%xhighEpoch AI
GSO25.5%highEpoch AI
GSO31.4%#10 of 31, top 33%xhighEpoch AI
WeirdML57.4%noneEpoch AI
WeirdML77.7%#13 of 119, top 11%xhighEpoch AI
LMArena Coding1497#24 of 294, top 9%highLMArena2026-10-08
MirrorCode15.6%#7 of 9, top 78%highEpoch AI2026-08-10
ALE-Bench1,607#13 of 105, top 13%highEpoch AI
ALE-Bench1,521mediumEpoch AI
ALE-Bench1,086noneEpoch AI
AlgoTune1.85#3 of 18, top 17%highEpoch AI

Agentic & Tool Use

GPT-5.4 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Terminal-Bench81.8%#2 of 41, top 5%Epoch AI
APEX-Agents52.4%#23 of 49, top 47%Epoch AI
τ²-bench Banking39.4%#10 of 26, top 39%xhighτ²-bench2026-03-25
DeepResearch Bench35.1%#24 of 24, top 100%lowEpoch AI
PostTrainBench19%#11 of 11, top 100%highEpoch AI
GBAEval45.1%#10 of 23, top 44%Epoch AI
LMArena Search1197#16 of 32, top 50%LMArena2026-08-24
METR Time Horizons74.3%#7 of 32, top 22%xhighEpoch AI
Vending-Bench 26,144#19 of 60, top 32%Epoch AI

Reasoning

GPT-5.4 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-267.5%highEpoch AI
ARC-AGI-229.2%lowEpoch AI
ARC-AGI-255.4%mediumEpoch AI
ARC-AGI-274%#18 of 83, top 22%xhighEpoch AI
Kagi LLM Benchmark63.8%#34 of 99, top 35%Kagi LLM Benchmark
NYT Connections (extended)91.3%#15 of 91, top 17%xhigh reasoningLech Mazur benchmarks
ARC-AGI-192.7%highEpoch AI
ARC-AGI-168.2%lowEpoch AI
ARC-AGI-186.2%mediumEpoch AI
ARC-AGI-193.7%#19 of 83, top 23%xhighEpoch AI
CritPt23.4%#18 of 134, top 14%xhighEpoch AI
Chess Puzzles38%highEpoch AI2026-07-15
Chess Puzzles20%lowEpoch AI2026-07-15
Chess Puzzles38%mediumEpoch AI2026-07-15
Chess Puzzles5%noneEpoch AI2026-07-15
Chess Puzzles44%#15 of 129, top 12%xhighEpoch AI2026-03-11
EnigmaEval16%#9 of 38, top 24%xhighEpoch AI
Thematic Generalization80%#2 of 23, top 9%xhigh reasoningLech Mazur benchmarks
EBR-Bench25.4%#12 of 24, top 50%xhighEpoch AI2026-06-25
LMArena Hard Prompts1485#24 of 297, top 9%highLMArena2026-10-08
Mystery Game Puzzles17%lowEpoch AI2026-08-27
Mystery Game Puzzles28%mediumEpoch AI2026-08-28
Mystery Game Puzzles16%noneEpoch AI2026-08-30
Mystery Game Puzzles37%#16 of 74, top 22%xhighEpoch AI2026-07-24
DTBench94.4%#22 of 151, top 15%xhighEpoch AI
LMCA52%#19 of 125, top 16%xhighEpoch AI
Epoch Capabilities Index156.81#17 of 213, top 8%Epoch AI2026-03-05
ForecastBench59.5#38 of 72, top 53%Epoch AI

Math

GPT-5.4 Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)78.6%#17 of 81, top 21%xhighEpoch AI2026-06-11
FrontierMath Tier 449%#18 of 63, top 29%xhighEpoch AI2026-06-11
MathArena Final-Answer Competitions83.1%#5 of 29, top 18%xhighMathArena
OTIS Mock AIME 2024-202597.8%#25 of 173, top 15%highEpoch AI2026-07-15
OTIS Mock AIME 2024-202584.4%lowEpoch AI2026-07-15
OTIS Mock AIME 2024-202595.6%mediumEpoch AI2026-07-15
OTIS Mock AIME 2024-202557.8%noneEpoch AI2026-07-15
OTIS Mock AIME 2024-202595.3%xhighEpoch AI2026-03-06
ProofBench56%#24 of 77, top 32%xhighEpoch AI
LMArena Math1488#21 of 285, top 8%highLMArena2026-10-08
FrontierMath (Feb 2025 set)47.6%#4 of 68, top 6%xhighEpoch AI2026-03-06
FrontierMath Tier 4 (v1)27.1%#7 of 55, top 13%xhighEpoch AI2026-03-06

Knowledge

GPT-5.4 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond89.9%highEpoch AI2026-07-15
GPQA Diamond84.8%lowEpoch AI2026-07-15
GPQA Diamond88.9%mediumEpoch AI2026-07-15
GPQA Diamond74.7%noneEpoch AI2026-07-15
GPQA Diamond93.3%#17 of 186, top 10%xhighEpoch AI2026-03-06
Humanity's Last Exam36.2%#8 of 41, top 20%xhighEpoch AI
SimpleQA Verified45.1%#36 of 77, top 47%xhighEpoch AI2026-08-27
Vectara Hallucination Rate (lower is better)7%#29 of 96, top 31%Vectara Hallucination Leaderboard
LMArena Expert1507#19 of 273, top 7%highLMArena2026-10-08

Multimodal

GPT-5.4 Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1303#15 of 122, top 13%highLMArena2026-10-09
Blueprint-Bench 227.1%#19 of 31, top 62%Epoch AI
Furniture Assembly37.5%#17 of 31, top 55%xhighEpoch AI2026-09-10
LMArena Document1471#11 of 38, top 29%LMArena2026-09-13

Multilingual

GPT-5.4 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1465#23 of 297, top 8%highLMArena2026-10-08
LMArena Chinese1519#28 of 285, top 10%highLMArena2026-10-08
LMArena French1493#17 of 223, top 8%highLMArena2026-10-08
LMArena German1472#23 of 231, top 10%highLMArena2026-10-08
LMArena Japanese1485#13 of 211, top 7%highLMArena2026-10-08
LMArena Korean1448#19 of 213, top 9%highLMArena2026-10-08
LMArena Russian1480#21 of 283, top 8%highLMArena2026-10-08
LMArena Spanish1454#46 of 226, top 21%highLMArena2026-10-08

Instruction Following

GPT-5.4 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1469#24 of 298, top 9%highLMArena2026-10-08

Long Context

GPT-5.4 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
CL-bench27.9%Best of 19xhighEpoch AI
CL-bench Life13.8%Epoch AI
CL-bench Life19.3%highEpoch AI
CL-bench Life21.7%#2 of 13, top 16%xhighEpoch AI
LMArena Longer Query1473#32 of 291, top 11%highLMArena2026-10-08

Writing & Preference

GPT-5.4 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1469#28 of 297, top 10%highLMArena2026-10-08
LMArena Creative Writing1439#40 of 295, top 14%highLMArena2026-10-08
EQ-Bench Creative Writing1840#20 of 115, top 18%EQ-Bench
EQ-Bench 41272#7 of 28, top 25%EQ-Bench
LMArena Multi-Turn1482#17 of 295, top 6%highLMArena2026-10-08

API pricing by provider

GPT-5.4 API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$2.50$15$0.252026-10-10
bedrock$2.75$16.50$0.282026-10-10
openai$2.50$15$0.252026-10-10
openrouter$2.50$15$0.252026-10-10

Compare GPT-5.4

Other OpenAI models

Frequently asked questions

How good is GPT-5.4?

GPT-5.4 by OpenAI ranks 16th of 354 ranked models on the Noometry Index as of October 2026, with a score of 59.4. Its strongest category is long context, where it ranks 8th. API pricing starts at $2.50 per million input tokens and $15 per million output tokens, with a 1.05M-token context window.

How much does GPT-5.4 cost?

GPT-5.4 costs $2.50 per million input tokens and $15 per million output tokens on OpenAI's own API, with cached input at $0.25.

What is GPT-5.4's context window?

GPT-5.4 accepts up to 1.05M tokens of input and can write up to 128K tokens in one response.

Is GPT-5.4 open source?

No. GPT-5.4 is proprietary and available only through OpenAI's API and partner platforms.

How fast is GPT-5.4?

GPT-5.4 generated about 12 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are GPT-5.4's strengths and weaknesses?

Relative to other ranked models, GPT-5.4 places best in long context, knowledge, reasoning and lowest in multimodal, coding, instruction following.

What is GPT-5.4 best at?

Its best category is long context, where it ranks 8th on Noometry.