Model comparison

GPT-4 vs GPT-5.6 Sol

GPT-5.6 Sol is the stronger model overall, scoring 65.0 to 29.1 on the Noometry Index.

Last verified . 26 shared benchmarks.

GPT-4 OpenAI

29.1

Rank #316 Confirmed

GPT-5.6 Sol OpenAI

65.0

Rank #7 Confirmed

Summary

  • They share 26 benchmarks with published results for both. GPT-4 scores higher in 0 categories and GPT-5.6 Sol in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-5.6 Sol leads 85.6 to 10.8.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 1.1% for GPT-4 and 100% for GPT-5.6 Sol.
  • GPT-5.6 Sol is cheaper at $4 / $20 per million input/output tokens, against $30 / $60 for GPT-4.
  • GPT-5.6 Sol accepts more context: 1.05M tokens versus 8K.

Side by side

GPT-4 and GPT-5.6 Sol specifications
GPT-4GPT-5.6 Sol
ProviderOpenAIOpenAI
Noometry Index29.165.0
Released2023-03-142026-07-09
WeightsProprietaryProprietary
Context window8K1.05M
Max output8K128K
Input $ / M tokens$30$4
Output $ / M tokens$60$20
Results tracked3865

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.6 Sol leads

GPT-4: 31.6 (#283), GPT-5.6 Sol: 65.1 (#7)

Coding benchmarks
BenchmarkGPT-4GPT-5.6 Sol
WeirdML12.4%89.4%
LMArena Coding12541498
DeepSWE—72.7%
FrontierCode—47.5%
CursorBench—41.7%
LMArena WebDev—1618
FrontierSWE—32.2%
SciCode—57.1%
GSO—76.5%
BigCodeBench Instruct46%—
MirrorCode—20%
BigCodeBench Complete57.2%—
ALE-Bench—2,177
HumanEval+79.3%—

Agentic & Tool Use Not comparable

GPT-4: —, GPT-5.6 Sol: 50.3 (#7)

Agentic & Tool Use benchmarks
BenchmarkGPT-4GPT-5.6 Sol
APEX-Agents—51.4%
OSWorld 2.0—27.3%
τ²-bench Banking—46.9%
PostTrainBench—36.2%
BALROG—60%
GBAEval—52.6%
GDP.pdf—30.7%
LMArena Search—1257
METR Time Horizons36.1%—
Vending-Bench 2—9,619

Reasoning GPT-5.6 Sol leads

GPT-4: 17.8 (#289), GPT-5.6 Sol: 74.8 (#8)

Reasoning benchmarks
BenchmarkGPT-4GPT-5.6 Sol
Chess Puzzles4%64%
LMArena Hard Prompts12411484
Mystery Game Puzzles12%58%
DTBench62.7%96%
LMCA17.1%59.2%
Epoch Capabilities Index125.89161.66
ARC-AGI-2—92.5%
SimpleBench—71.7%
Kagi LLM Benchmark—67%
NYT Connections (extended)—93.8%
ARC-AGI-1—97.5%
CritPt—32.3%
EnigmaEval—37.1%
EBR-Bench—44.8%
Surface Evolver Bench—93.1%
Bench to the Future 3—0.14
BIG-Bench Hard75.1%—
ForecastBench57.8—
HellaSwag95.3%—
WinoGrande87.5%—

Math GPT-5.6 Sol leads

GPT-4: 10.8 (#309), GPT-5.6 Sol: 85.6 (#9)

Math benchmarks
BenchmarkGPT-4GPT-5.6 Sol
OTIS Mock AIME 2024-20251.1%100%
LMArena Math12691474
FrontierMath (Tiers 1-3)—89.1%
FrontierMath Tier 4—82.9%
ProofBench—83%
MATH Level 523%—
FrontierMath Erdős—0%
GSM8K92%—

Knowledge GPT-5.6 Sol leads

GPT-4: 18.4 (#282), GPT-5.6 Sol: 64.3 (#18)

Knowledge benchmarks
BenchmarkGPT-4GPT-5.6 Sol
GPQA Diamond35.7%93.5%
LMArena Expert12111516
SimpleQA Verified—69.7%
Vectara Hallucination Rate—12.4%
MMLU86.4%—
TriviaQA84.8%—

Multimodal Not comparable

GPT-4: —, GPT-5.6 Sol: 48.6 (#9)

Multimodal benchmarks
BenchmarkGPT-4GPT-5.6 Sol
LMArena Vision—1281
Blueprint-Bench 2—33.6%
Furniture Assembly—56.7%
LMArena Document—1483

Multilingual GPT-5.6 Sol leads

GPT-4: 40.6 (#215), GPT-5.6 Sol: 55.3 (#32)

Multilingual benchmarks
BenchmarkGPT-4GPT-5.6 Sol
LMArena Non-English12461452
LMArena Chinese12421527
LMArena French12831477
LMArena German12511476
LMArena Japanese12091471
LMArena Korean11841442
LMArena Russian12511468
LMArena Spanish12611441

Instruction Following GPT-5.6 Sol leads

GPT-4: 65.3 (#222), GPT-5.6 Sol: 77.7 (#16)

Instruction Following benchmarks
BenchmarkGPT-4GPT-5.6 Sol
LMArena Instruction Following12411482

Long Context GPT-5.6 Sol leads

GPT-4: 37.7 (#212), GPT-5.6 Sol: 45.4 (#42)

Long Context benchmarks
BenchmarkGPT-4GPT-5.6 Sol
LMArena Longer Query12441480

Writing & Preference GPT-5.6 Sol leads

GPT-4: 34.9 (#268), GPT-5.6 Sol: 73.3 (#12)

Writing & Preference benchmarks
BenchmarkGPT-4GPT-5.6 Sol
LMArena Text12631457
LMArena Creative Writing12441448
EQ-Bench Creative Writing7521972
LMArena Multi-Turn12571460
EQ-Bench 4—1250

Frequently asked questions

Is GPT-4 better than GPT-5.6 Sol?

GPT-5.6 Sol is the stronger model overall, scoring 65.0 to 29.1 on the Noometry Index.

Which is cheaper, GPT-4 or GPT-5.6 Sol?

GPT-5.6 Sol is cheaper. It lists at $4 per million input tokens and $20 per million output tokens; GPT-4 lists at $30 and $60.

Is GPT-4 or GPT-5.6 Sol better for coding?

GPT-5.6 Sol scores higher on coding benchmarks: 65.1 versus 31.6 in the Noometry coding category.

Which has the bigger context window?

GPT-5.6 Sol does, with 1.05M tokens against 8K.

How many benchmarks do GPT-4 and GPT-5.6 Sol share?

26 benchmarks have published results for both models. GPT-4 has 38 scored results on Noometry and GPT-5.6 Sol has 65.

Related comparisons

Go deeper