Model comparison

Claude 3 Opus vs GPT-5.6 Sol

GPT-5.6 Sol is the stronger model overall, scoring 65.0 to 29.5 on the Noometry Index.

Last verified . 28 shared benchmarks.

Claude 3 Opus Anthropic

29.5

Rank #310 Confirmed

GPT-5.6 Sol OpenAI

65.0

Rank #7 Confirmed

Summary

  • They share 28 benchmarks with published results for both. Claude 3 Opus scores higher in 0 categories and GPT-5.6 Sol in 10 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-5.6 Sol leads 85.6 to 14.8.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 4.7% for Claude 3 Opus and 100% for GPT-5.6 Sol.

Side by side

Claude 3 Opus and GPT-5.6 Sol specifications
Claude 3 OpusGPT-5.6 Sol
ProviderAnthropicOpenAI
Noometry Index29.565.0
Released2024-02-292026-07-09
WeightsProprietaryProprietary
Context window—1.05M
Max output—128K
Input $ / M tokens—$4
Output $ / M tokens—$20
Results tracked4665

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.6 Sol leads

Claude 3 Opus: 32.9 (#267), GPT-5.6 Sol: 65.1 (#7)

Coding benchmarks
BenchmarkClaude 3 OpusGPT-5.6 Sol
WeirdML19.2%89.4%
LMArena Coding12641498
DeepSWE—72.7%
FrontierCode—47.5%
CursorBench—41.7%
LMArena WebDev—1618
FrontierSWE—32.2%
SciCode—57.1%
GSO—76.5%
BigCodeBench Instruct45.5%—
LiveBench Coding38.6%—
MirrorCode—20%
BigCodeBench Complete57.4%—
ALE-Bench—2,177
HumanEval+77.4%—
MBPP+73.3%—

Agentic & Tool Use GPT-5.6 Sol leads

Claude 3 Opus: 24.6 (#116), GPT-5.6 Sol: 50.3 (#7)

Agentic & Tool Use benchmarks
BenchmarkClaude 3 OpusGPT-5.6 Sol
APEX-Agents—51.4%
OSWorld 2.0—27.3%
τ²-bench Banking—46.9%
Cybench10%—
PostTrainBench—36.2%
BALROG—60%
GBAEval—52.6%
GDP.pdf—30.7%
LMArena Search—1257
METR Time Horizons29.5%—
Vending-Bench 2—9,619

Reasoning GPT-5.6 Sol leads

Claude 3 Opus: 14.6 (#324), GPT-5.6 Sol: 74.8 (#8)

Reasoning benchmarks
BenchmarkClaude 3 OpusGPT-5.6 Sol
SimpleBench23.5%71.7%
Chess Puzzles5%64%
EnigmaEval0.8%37.1%
LMArena Hard Prompts12451484
DTBench61.6%96%
LMCA17%59.2%
Epoch Capabilities Index126.91161.66
ARC-AGI-2—92.5%
Kagi LLM Benchmark—67%
NYT Connections (extended)—93.8%
ARC-AGI-1—97.5%
CritPt—32.3%
EBR-Bench—44.8%
LiveBench Reasoning40.6%—
Mystery Game Puzzles—58%
LiveBench Data Analysis57.9%—
Surface Evolver Bench—93.1%
Bench to the Future 3—0.14
ForecastBench58.4—
LiveBench49.2%—
WinoGrande88.5%—

Math GPT-5.6 Sol leads

Claude 3 Opus: 14.8 (#299), GPT-5.6 Sol: 85.6 (#9)

Math benchmarks
BenchmarkClaude 3 OpusGPT-5.6 Sol
OTIS Mock AIME 2024-20254.7%100%
LMArena Math12731474
FrontierMath (Tiers 1-3)—89.1%
FrontierMath Tier 4—82.9%
ProofBench—83%
LiveBench Math43.6%—
MATH Level 537.5%—
FrontierMath Erdős—0%

Knowledge GPT-5.6 Sol leads

Claude 3 Opus: 24.5 (#267), GPT-5.6 Sol: 64.3 (#18)

Knowledge benchmarks
BenchmarkClaude 3 OpusGPT-5.6 Sol
GPQA Diamond47.2%93.5%
SimpleQA Verified12.6%69.7%
LMArena Expert12231516
Confabulations22.7%—
Vectara Hallucination Rate—12.4%
MMLU84.6%—

Multimodal GPT-5.6 Sol leads

Claude 3 Opus: 27.1 (#116), GPT-5.6 Sol: 48.6 (#9)

Multimodal benchmarks
BenchmarkClaude 3 OpusGPT-5.6 Sol
LMArena Vision10231281
Blueprint-Bench 2—33.6%
Furniture Assembly—56.7%
LMArena Document—1483

Multilingual GPT-5.6 Sol leads

Claude 3 Opus: 41.4 (#207), GPT-5.6 Sol: 55.3 (#32)

Multilingual benchmarks
BenchmarkClaude 3 OpusGPT-5.6 Sol
LMArena Non-English12581452
LMArena Chinese12481527
LMArena French12751477
LMArena German12581476
LMArena Japanese12041471
LMArena Korean11871442
LMArena Russian12801468
LMArena Spanish12461441

Instruction Following GPT-5.6 Sol leads

Claude 3 Opus: 64.1 (#228), GPT-5.6 Sol: 77.7 (#16)

Instruction Following benchmarks
BenchmarkClaude 3 OpusGPT-5.6 Sol
LMArena Instruction Following12481482
LiveBench Instruction Following63.9%—

Long Context GPT-5.6 Sol leads

Claude 3 Opus: 38.2 (#202), GPT-5.6 Sol: 45.4 (#42)

Long Context benchmarks
BenchmarkClaude 3 OpusGPT-5.6 Sol
LMArena Longer Query12591480

Writing & Preference GPT-5.6 Sol leads

Claude 3 Opus: 47.2 (#213), GPT-5.6 Sol: 73.3 (#12)

Writing & Preference benchmarks
BenchmarkClaude 3 OpusGPT-5.6 Sol
LMArena Text12621457
LMArena Creative Writing12351448
LMArena Multi-Turn12751460
EQ-Bench Creative Writing—1972
EQ-Bench 4—1250
LiveBench Language50.4%—

Frequently asked questions

Is Claude 3 Opus better than GPT-5.6 Sol?

GPT-5.6 Sol is the stronger model overall, scoring 65.0 to 29.5 on the Noometry Index.

Is Claude 3 Opus or GPT-5.6 Sol better for coding?

GPT-5.6 Sol scores higher on coding benchmarks: 65.1 versus 32.9 in the Noometry coding category.

How many benchmarks do Claude 3 Opus and GPT-5.6 Sol share?

28 benchmarks have published results for both models. Claude 3 Opus has 46 scored results on Noometry and GPT-5.6 Sol has 65.

Related comparisons

Go deeper