Model comparison

GPT-4 vs GPT-5.5 Instant

GPT-5.5 Instant is the stronger model overall, scoring 42.7 to 29.1 on the Noometry Index.

Last verified . 21 shared benchmarks.

GPT-4 OpenAI

29.1

Rank #316 Confirmed

GPT-5.5 Instant OpenAI

42.7

Rank #110 Confirmed

Summary

  • They share 21 benchmarks with published results for both. GPT-4 scores higher in 0 categories and GPT-5.5 Instant in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where GPT-5.5 Instant leads 48.9 to 18.4.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 1.1% for GPT-4 and 68.1% for GPT-5.5 Instant.

Side by side

GPT-4 and GPT-5.5 Instant specifications
GPT-4GPT-5.5 Instant
ProviderOpenAIOpenAI
Noometry Index29.142.7
Released2023-03-142026-05-05
WeightsProprietaryProprietary
Context window8K—
Max output8K—
Input $ / M tokens$30—
Output $ / M tokens$60—
Results tracked3827

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.5 Instant leads

GPT-4: 31.6 (#283), GPT-5.5 Instant: 44.3 (#74)

Coding benchmarks
BenchmarkGPT-4GPT-5.5 Instant
LMArena Coding12541433
SciCode—48.6%
WeirdML12.4%—
BigCodeBench Instruct46%—
BigCodeBench Complete57.2%—
HumanEval+79.3%—

Agentic & Tool Use Not comparable

GPT-4: —, GPT-5.5 Instant: —

Agentic & Tool Use benchmarks
BenchmarkGPT-4GPT-5.5 Instant
METR Time Horizons36.1%—

Reasoning GPT-5.5 Instant leads

GPT-4: 17.8 (#289), GPT-5.5 Instant: 24.9 (#155)

Reasoning benchmarks
BenchmarkGPT-4GPT-5.5 Instant
Chess Puzzles4%12%
LMArena Hard Prompts12411426
Epoch Capabilities Index125.89142.52
CritPt—0%
Mystery Game Puzzles12%—
DTBench62.7%—
LMCA17.1%—
BIG-Bench Hard75.1%—
ForecastBench57.8—
HellaSwag95.3%—
WinoGrande87.5%—

Math GPT-5.5 Instant leads

GPT-4: 10.8 (#309), GPT-5.5 Instant: 26.5 (#259)

Math benchmarks
BenchmarkGPT-4GPT-5.5 Instant
OTIS Mock AIME 2024-20251.1%68.1%
LMArena Math12691420
FrontierMath (Tiers 1-3)—26.3%
FrontierMath Tier 4—2.4%
MATH Level 523%—
GSM8K92%—

Knowledge GPT-5.5 Instant leads

GPT-4: 18.4 (#282), GPT-5.5 Instant: 48.9 (#74)

Knowledge benchmarks
BenchmarkGPT-4GPT-5.5 Instant
GPQA Diamond35.7%82.5%
LMArena Expert12111409
MMLU86.4%—
TriviaQA84.8%—

Multimodal Not comparable

GPT-4: —, GPT-5.5 Instant: 40.0 (#52)

Multimodal benchmarks
BenchmarkGPT-4GPT-5.5 Instant
LMArena Vision—1250
LMArena Document—1403

Multilingual GPT-5.5 Instant leads

GPT-4: 40.6 (#215), GPT-5.5 Instant: 52.8 (#80)

Multilingual benchmarks
BenchmarkGPT-4GPT-5.5 Instant
LMArena Non-English12461417
LMArena Chinese12421456
LMArena French12831428
LMArena German12511411
LMArena Japanese12091408
LMArena Korean11841392
LMArena Russian12511431
LMArena Spanish12611429

Instruction Following GPT-5.5 Instant leads

GPT-4: 65.3 (#222), GPT-5.5 Instant: 74.2 (#100)

Instruction Following benchmarks
BenchmarkGPT-4GPT-5.5 Instant
LMArena Instruction Following12411406

Long Context GPT-5.5 Instant leads

GPT-4: 37.7 (#212), GPT-5.5 Instant: 43.4 (#96)

Long Context benchmarks
BenchmarkGPT-4GPT-5.5 Instant
LMArena Longer Query12441422

Writing & Preference GPT-5.5 Instant leads

GPT-4: 34.9 (#268), GPT-5.5 Instant: 61.8 (#85)

Writing & Preference benchmarks
BenchmarkGPT-4GPT-5.5 Instant
LMArena Text12631419
LMArena Creative Writing12441419
LMArena Multi-Turn12571433
EQ-Bench Creative Writing752—

Frequently asked questions

Is GPT-4 better than GPT-5.5 Instant?

GPT-5.5 Instant is the stronger model overall, scoring 42.7 to 29.1 on the Noometry Index.

Is GPT-4 or GPT-5.5 Instant better for coding?

GPT-5.5 Instant scores higher on coding benchmarks: 44.3 versus 31.6 in the Noometry coding category.

How many benchmarks do GPT-4 and GPT-5.5 Instant share?

21 benchmarks have published results for both models. GPT-4 has 38 scored results on Noometry and GPT-5.5 Instant has 27.

Related comparisons

Go deeper