Model comparison

GPT-5 vs GPT-5.5

GPT-5.5 is the stronger model overall, scoring 63.4 to 50.9 on the Noometry Index. GPT-5 costs 3.3× less per token, which makes it the better buy when GPT-5.5's lead doesn't matter for your workload.

Last verified . 50 shared benchmarks.

GPT-5 OpenAI

50.9

Rank #45 Confirmed

GPT-5.5 OpenAI

63.4

Rank #9 Confirmed

Summary

  • They share 50 benchmarks with published results for both. GPT-5 scores higher in 1 category and GPT-5.5 in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where GPT-5.5 leads 72.8 to 38.3.
  • The biggest single-benchmark swing is ARC-AGI-2: 9.9% for GPT-5 and 85% for GPT-5.5.
  • GPT-5 is cheaper at $1.25 / $10 per million input/output tokens, against $5 / $30 for GPT-5.5.
  • GPT-5.5 accepts more context: 1.05M tokens versus 400K.

Side by side

GPT-5 and GPT-5.5 specifications
GPT-5GPT-5.5
ProviderOpenAIOpenAI
Noometry Index50.963.4
Released2025-08-072026-04-23
WeightsProprietaryProprietary
Context window400K1.05M
Max output128K128K
Input $ / M tokens$1.25$5
Output $ / M tokens$10$30
Results tracked6971

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.5 leads

GPT-5: 50.3 (#47), GPT-5.5: 58.2 (#17)

Coding benchmarks
BenchmarkGPT-5GPT-5.5
SWE-bench Verified73.6%80.6%
LMArena WebDev14181513
SciCode42.9%56.1%
GSO6.9%40.2%
WeirdML60.7%84.9%
LMArena Coding14361494
ALE-Bench1,1621,943
DeepSWE—67%
FrontierCode—43%
SWE-bench Verified (bash only)65%—
Aider Polyglot88%—
MirrorCode—10%
AlgoTune1.67—

Agentic & Tool Use GPT-5.5 leads

GPT-5: 33.1 (#56), GPT-5.5: 50.7 (#6)

Agentic & Tool Use benchmarks
BenchmarkGPT-5GPT-5.5
Terminal-Bench49.6%84.7%
Remote Labor Index1.7%6.3%
DeepResearch Bench49.6%54%
LMArena Search11331242
APEX-Agents—55.1%
OSWorld 2.0—13%
GDPval34.8%—
τ²-bench Banking—44.6%
PostTrainBench—27.2%
BALROG32.8%—
ExploitBench—47.4%
GBAEval—53.2%
GDP.pdf—26%
METR Time Horizons69.6%—
Vending-Bench 2—7,524

Reasoning GPT-5.5 leads

GPT-5: 38.3 (#64), GPT-5.5: 72.8 (#11)

Reasoning benchmarks
BenchmarkGPT-5GPT-5.5
ARC-AGI-29.9%85%
SimpleBench56.7%69%
Kagi LLM Benchmark72.7%88.8%
ARC-AGI-165.7%95%
CritPt12.6%27.1%
Chess Puzzles37%54%
EBR-Bench12.7%34.3%
LMArena Hard Prompts14161489
Mystery Game Puzzles23%56%
DTBench90.7%96%
LMCA40%54.3%
Epoch Capabilities Index150159.1
ForecastBench61.460.6
NYT Connections (extended)—96.2%
EnigmaEval10.5%—
Surface Evolver Bench—88.1%
Bench to the Future 3—0.14

Math GPT-5.5 leads

GPT-5: 55.0 (#44), GPT-5.5: 81.7 (#11)

Knowledge GPT-5.5 leads

GPT-5: 56.6 (#43), GPT-5.5: 64.4 (#17)

Knowledge benchmarks
BenchmarkGPT-5GPT-5.5
GPQA Diamond86.2%94%
SimpleQA Verified50.1%63%
Vectara Hallucination Rate14.7%9.3%
LMArena Expert14191508
Humanity's Last Exam25.3%—
MMLU-Pro86.3%—
Confabulations10.3%—
GPQA (HELM)79.2%—

Multimodal Too close to call

GPT-5: 46.8 (#13), GPT-5.5: 46.9 (#12)

Multimodal benchmarks
BenchmarkGPT-5GPT-5.5
LMArena Vision12321297
GeoBench81%—
VPCT66%—
Blueprint-Bench 2—36.2%
Furniture Assembly—44.2%
LMArena Document—1486

Multilingual GPT-5.5 leads

GPT-5: 51.4 (#110), GPT-5.5: 56.4 (#20)

Multilingual benchmarks
BenchmarkGPT-5GPT-5.5
LMArena Non-English13971467
LMArena Chinese14221533
LMArena French14101486
LMArena German14161480
LMArena Japanese14091498
LMArena Korean13601460
LMArena Russian14061473
LMArena Spanish13991468

Instruction Following GPT-5.5 leads

GPT-5: 73.8 (#113), GPT-5.5: 77.5 (#18)

Instruction Following benchmarks
BenchmarkGPT-5GPT-5.5
LMArena Instruction Following13881479
IFEval87.5%—

Long Context GPT-5 leads

GPT-5: 69.5 (#2), GPT-5.5: 48.3 (#12)

Long Context benchmarks
BenchmarkGPT-5GPT-5.5
LMArena Longer Query13991484
Fiction.LiveBench97.2%—
CL-bench Life—22.2%

Writing & Preference GPT-5.5 leads

GPT-5: 63.4 (#65), GPT-5.5: 72.7 (#13)

Writing & Preference benchmarks
BenchmarkGPT-5GPT-5.5
LMArena Text14061472
LMArena Creative Writing13651455
EQ-Bench Creative Writing16271844
LMArena Multi-Turn14261476
Short-Story Creative Writing86%—
WildBench85.7%—
EQ-Bench 4—1315

Frequently asked questions

Is GPT-5 better than GPT-5.5?

GPT-5.5 is the stronger model overall, scoring 63.4 to 50.9 on the Noometry Index. GPT-5 costs 3.3× less per token, which makes it the better buy when GPT-5.5's lead doesn't matter for your workload.

Which is cheaper, GPT-5 or GPT-5.5?

GPT-5 is cheaper. It lists at $1.25 per million input tokens and $10 per million output tokens; GPT-5.5 lists at $5 and $30.

Is GPT-5 or GPT-5.5 better for coding?

GPT-5.5 scores higher on coding benchmarks: 58.2 versus 50.3 in the Noometry coding category.

Which has the bigger context window?

GPT-5.5 does, with 1.05M tokens against 400K.

How many benchmarks do GPT-5 and GPT-5.5 share?

50 benchmarks have published results for both models. GPT-5 has 69 scored results on Noometry and GPT-5.5 has 71.

Related comparisons

Go deeper