Model comparison

Claude Sonnet 4 vs Claude Sonnet 5.5

Claude Sonnet 5.5 is the stronger model overall, scoring 61.9 to 40.8 on the Noometry Index.

Last verified . 19 shared benchmarks.

Claude Sonnet 4 Anthropic

40.8

Rank #145 Confirmed

Claude Sonnet 5.5 Anthropic

61.9

Rank #10 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Claude Sonnet 4 scores higher in 0 categories and Claude Sonnet 5.5 in 10 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Sonnet 5.5 leads 87.9 to 43.3.
  • The biggest single-benchmark swing is CritPt: 0.3% for Claude Sonnet 4 and 31.4% for Claude Sonnet 5.5.
  • Claude Sonnet 5.5 is cheaper at $2 / $10 per million input/output tokens, against $3 / $15 for Claude Sonnet 4.
  • Claude Sonnet 5.5 accepts more context: 1M tokens versus 200K.

Side by side

Claude Sonnet 4 and Claude Sonnet 5.5 specifications
Claude Sonnet 4Claude Sonnet 5.5
ProviderAnthropicAnthropic
Noometry Index40.861.9
Released2025-05-222026-09-28
WeightsProprietaryProprietary
Context window200K1M
Max output64K128K
Input $ / M tokens$3$2
Output $ / M tokens$15$10
Results tracked5832

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Sonnet 5.5 leads

Claude Sonnet 4: 43.5 (#88), Claude Sonnet 5.5: 67.3 (#6)

Coding benchmarks
BenchmarkClaude Sonnet 4Claude Sonnet 5.5
SciCode40%61%
LMArena Coding14141513
ALE-Bench655.351,819
FrontierCode—52.1%
SWE-bench Verified (bash only)64.9%—
Aider Polyglot61.3%—
CursorBench—55.5%
LMArena WebDev—1774
FrontierSWE—61.9%
GSO4.9%—
WeirdML46.1%—

Agentic & Tool Use Claude Sonnet 5.5 leads

Claude Sonnet 4: 38.5 (#31), Claude Sonnet 5.5: 45.0 (#16)

Agentic & Tool Use benchmarks
BenchmarkClaude Sonnet 4Claude Sonnet 5.5
APEX-Agents—75.5%
TheAgentCompany33.1%—
Cybench35%—
DeepResearch Bench46.6%—
OSWorld43.9%—
METR Time Horizons62%—

Reasoning Claude Sonnet 5.5 leads

Claude Sonnet 4: 22.9 (#187), Claude Sonnet 5.5: 54.0 (#28)

Reasoning benchmarks
BenchmarkClaude Sonnet 4Claude Sonnet 5.5
CritPt0.3%31.4%
LMArena Hard Prompts13721495
Epoch Capabilities Index141.69165.03
ARC-AGI-25.9%—
SimpleBench45.5%—
Kagi LLM Benchmark73%—
NYT Connections (extended)—80.5%
ARC-AGI-140%—
EnigmaEval3.1%—
Mystery Game Puzzles—65%
DTBench77.1%—
LMCA29%—
ForecastBench60.2—

Math Claude Sonnet 5.5 leads

Claude Sonnet 4: 43.3 (#80), Claude Sonnet 5.5: 87.9 (#6)

Math benchmarks
BenchmarkClaude Sonnet 4Claude Sonnet 5.5
OTIS Mock AIME 2024-202571.1%100%
LMArena Math13751510
FrontierMath (Tiers 1-3)—88.8%
FrontierMath Tier 4—80.5%
ProofBench—100%
Omni-MATH60.2%—
MATH Level 584.4%—
FrontierMath (Feb 2025 set)4.1%—
FrontierMath Erdős—2.9%
FrontierMath Tier 4 (v1)0%—

Knowledge Claude Sonnet 5.5 leads

Claude Sonnet 4: 41.8 (#108), Claude Sonnet 5.5: 66.0 (#12)

Knowledge benchmarks
BenchmarkClaude Sonnet 4Claude Sonnet 5.5
GPQA Diamond79.2%95.6%
LMArena Expert13721540
Humanity's Last Exam7.8%—
SimpleQA Verified—46.5%
MMLU-Pro84.3%—
Confabulations13.2%—
Vectara Hallucination Rate10.3%—
GPQA (HELM)70.6%—

Multimodal Claude Sonnet 5.5 leads

Claude Sonnet 4: 26.2 (#121), Claude Sonnet 5.5: 51.5 (#6)

Multimodal benchmarks
BenchmarkClaude Sonnet 4Claude Sonnet 5.5
LMArena Vision11911289
GeoBench37%—
VPCT34%—
Furniture Assembly—75%
MindCube44.8%—

Multilingual Claude Sonnet 5.5 leads

Claude Sonnet 4: 46.7 (#156), Claude Sonnet 5.5: 55.3 (#30)

Multilingual benchmarks
BenchmarkClaude Sonnet 4Claude Sonnet 5.5
LMArena Non-English13331452
LMArena Chinese13501522
LMArena Russian13551451
LMArena French1363—
LMArena German1331—
LMArena Japanese1302—
LMArena Korean1291—
LMArena Spanish1357—

Instruction Following Claude Sonnet 5.5 leads

Claude Sonnet 4: 71.7 (#145), Claude Sonnet 5.5: 78.3 (#11)

Instruction Following benchmarks
BenchmarkClaude Sonnet 4Claude Sonnet 5.5
LMArena Instruction Following13761495
IFEval84%—

Long Context Claude Sonnet 5.5 leads

Claude Sonnet 4: 33.7 (#259), Claude Sonnet 5.5: 45.9 (#28)

Long Context benchmarks
BenchmarkClaude Sonnet 4Claude Sonnet 5.5
LMArena Longer Query13981498
Fiction.LiveBench46.9%—

Writing & Preference Claude Sonnet 5.5 leads

Claude Sonnet 4: 57.1 (#132), Claude Sonnet 5.5: 66.0 (#40)

Writing & Preference benchmarks
BenchmarkClaude Sonnet 4Claude Sonnet 5.5
LMArena Text13511471
LMArena Creative Writing13451465
LMArena Multi-Turn13761474
Short-Story Creative Writing81.4%—
EQ-Bench Creative Writing1483—
WildBench83.8%—

Frequently asked questions

Is Claude Sonnet 4 better than Claude Sonnet 5.5?

Claude Sonnet 5.5 is the stronger model overall, scoring 61.9 to 40.8 on the Noometry Index.

Which is cheaper, Claude Sonnet 4 or Claude Sonnet 5.5?

Claude Sonnet 5.5 is cheaper. It lists at $2 per million input tokens and $10 per million output tokens; Claude Sonnet 4 lists at $3 and $15.

Is Claude Sonnet 4 or Claude Sonnet 5.5 better for coding?

Claude Sonnet 5.5 scores higher on coding benchmarks: 67.3 versus 43.5 in the Noometry coding category.

Which has the bigger context window?

Claude Sonnet 5.5 does, with 1M tokens against 200K.

How many benchmarks do Claude Sonnet 4 and Claude Sonnet 5.5 share?

19 benchmarks have published results for both models. Claude Sonnet 4 has 58 scored results on Noometry and Claude Sonnet 5.5 has 32.

Related comparisons

Go deeper