Model comparison

Claude Sonnet 5.5 vs GPT-4.5

Claude Sonnet 5.5 is the stronger model overall, scoring 61.9 to 37.2 on the Noometry Index.

Last verified . 16 shared benchmarks.

Claude Sonnet 5.5 Anthropic

61.9

Rank #10 Confirmed

GPT-4.5 OpenAI

37.2

Rank #208 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Claude Sonnet 5.5 scores higher in 10 categories and GPT-4.5 in 0 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Sonnet 5.5 leads 87.9 to 32.6.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 100% for Claude Sonnet 5.5 and 37.8% for GPT-4.5.

Side by side

Claude Sonnet 5.5 and GPT-4.5 specifications
Claude Sonnet 5.5GPT-4.5
ProviderAnthropicOpenAI
Noometry Index61.937.2
Released2026-09-282025-02-27
WeightsProprietaryProprietary
Context window1M—
Max output128K—
Input $ / M tokens$2—
Output $ / M tokens$10—
Results tracked3242

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Sonnet 5.5 leads

Claude Sonnet 5.5: 67.3 (#6), GPT-4.5: 42.2 (#109)

Coding benchmarks
BenchmarkClaude Sonnet 5.5GPT-4.5
LMArena Coding15131396
FrontierCode52.1%—
Aider Polyglot—44.9%
CursorBench55.5%—
LMArena WebDev1774—
FrontierSWE61.9%—
SciCode61%—
WeirdML—39.4%
LiveBench Coding—75.2%
ALE-Bench1,819—

Agentic & Tool Use Claude Sonnet 5.5 leads

Claude Sonnet 5.5: 45.0 (#16), GPT-4.5: 27.9 (#97)

Agentic & Tool Use benchmarks
BenchmarkClaude Sonnet 5.5GPT-4.5
APEX-Agents75.5%—
Cybench—17.5%

Reasoning Claude Sonnet 5.5 leads

Claude Sonnet 5.5: 54.0 (#28), GPT-4.5: 13.9 (#330)

Reasoning benchmarks
BenchmarkClaude Sonnet 5.5GPT-4.5
LMArena Hard Prompts14951403
Epoch Capabilities Index165.03136.74
ARC-AGI-2—0.8%
SimpleBench—34.5%
NYT Connections (extended)80.5%—
ARC-AGI-1—10.3%
CritPt31.4%—
EnigmaEval—3.2%
LiveBench Reasoning—71.1%
Mystery Game Puzzles65%—
LiveBench Data Analysis—64.3%
ForecastBench—61.7
LiveBench—69%

Math Claude Sonnet 5.5 leads

Claude Sonnet 5.5: 87.9 (#6), GPT-4.5: 32.6 (#211)

Math benchmarks
BenchmarkClaude Sonnet 5.5GPT-4.5
OTIS Mock AIME 2024-2025100%37.8%
LMArena Math15101412
FrontierMath (Tiers 1-3)88.8%—
FrontierMath Tier 480.5%—
ProofBench100%—
LiveBench Math—69.3%
MATH Level 5—78.6%
FrontierMath Erdős2.9%—

Knowledge Claude Sonnet 5.5 leads

Claude Sonnet 5.5: 66.0 (#12), GPT-4.5: 32.5 (#211)

Knowledge benchmarks
BenchmarkClaude Sonnet 5.5GPT-4.5
GPQA Diamond95.6%68.7%
LMArena Expert15401394
Humanity's Last Exam—5.4%
SimpleQA Verified46.5%—
Confabulations—13.6%

Multimodal Claude Sonnet 5.5 leads

Claude Sonnet 5.5: 51.5 (#6), GPT-4.5: 37.6 (#71)

Multimodal benchmarks
BenchmarkClaude Sonnet 5.5GPT-4.5
LMArena Vision12891195
VPCT—45%
Furniture Assembly75%—

Multilingual Claude Sonnet 5.5 leads

Claude Sonnet 5.5: 55.3 (#30), GPT-4.5: 52.5 (#83)

Multilingual benchmarks
BenchmarkClaude Sonnet 5.5GPT-4.5
LMArena Non-English14521413
LMArena Chinese15221421
LMArena Russian14511419
LMArena French—1418
LMArena German—1457
LMArena Japanese—1416
LMArena Korean—1392

Instruction Following Claude Sonnet 5.5 leads

Claude Sonnet 5.5: 78.3 (#11), GPT-4.5: 72.6 (#134)

Instruction Following benchmarks
BenchmarkClaude Sonnet 5.5GPT-4.5
LMArena Instruction Following14951404
LiveBench Instruction Following—72.3%

Long Context Claude Sonnet 5.5 leads

Claude Sonnet 5.5: 45.9 (#28), GPT-4.5: 40.4 (#155)

Long Context benchmarks
BenchmarkClaude Sonnet 5.5GPT-4.5
LMArena Longer Query14981406
Fiction.LiveBench—63.9%

Writing & Preference Claude Sonnet 5.5 leads

Claude Sonnet 5.5: 66.0 (#40), GPT-4.5: 56.9 (#134)

Writing & Preference benchmarks
BenchmarkClaude Sonnet 5.5GPT-4.5
LMArena Text14711417
LMArena Creative Writing14651394
LMArena Multi-Turn14741444
Short-Story Creative Writing—75.6%
EQ-Bench Creative Writing—1258
LiveBench Language—61.5%

Frequently asked questions

Is Claude Sonnet 5.5 better than GPT-4.5?

Claude Sonnet 5.5 is the stronger model overall, scoring 61.9 to 37.2 on the Noometry Index.

Is Claude Sonnet 5.5 or GPT-4.5 better for coding?

Claude Sonnet 5.5 scores higher on coding benchmarks: 67.3 versus 42.2 in the Noometry coding category.

How many benchmarks do Claude Sonnet 5.5 and GPT-4.5 share?

16 benchmarks have published results for both models. Claude Sonnet 5.5 has 32 scored results on Noometry and GPT-4.5 has 42.

Related comparisons

Go deeper