Model comparison

Claude 3.5 Sonnet vs Claude Sonnet 4.6

Claude Sonnet 4.6 is the stronger model overall, scoring 50.3 to 34.6 on the Noometry Index.

Last verified . 27 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

Claude Sonnet 4.6 Anthropic

50.3

Rank #50 Confirmed

Summary

  • They share 27 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 0 categories and Claude Sonnet 4.6 in 10 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Sonnet 4.6 leads 52.9 to 19.2.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 8.5% for Claude 3.5 Sonnet and 85.8% for Claude Sonnet 4.6.

Side by side

Claude 3.5 Sonnet and Claude Sonnet 4.6 specifications
Claude 3.5 SonnetClaude Sonnet 4.6
ProviderAnthropicAnthropic
Noometry Index34.650.3
Released2024-06-202026-02-17
WeightsProprietaryProprietary
Context window—1M
Max output—128K
Input $ / M tokens—$3
Output $ / M tokens—$15
Results tracked6057

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Sonnet 4.6 leads

Claude 3.5 Sonnet: 39.0 (#165), Claude Sonnet 4.6: 46.3 (#67)

Coding benchmarks
BenchmarkClaude 3.5 SonnetClaude Sonnet 4.6
WeirdML40%66.1%
LMArena Coding13421504
SWE-bench Verified—75.2%
DeepSWE—29.9%
FrontierCode—24.3%
Aider Polyglot51.6%—
LMArena WebDev—1522
SciCode—46.8%
GSO4.6%—
BigCodeBench Instruct46.8%—
LiveBench Coding67.1%—
BigCodeBench Complete58.6%—
CadEval48%—
ALE-Bench—1,327
HumanEval+81.7%—
MBPP+74.3%—

Agentic & Tool Use Claude Sonnet 4.6 leads

Claude 3.5 Sonnet: 32.3 (#67), Claude Sonnet 4.6: 39.1 (#28)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 SonnetClaude Sonnet 4.6
Terminal-Bench—53.4%
APEX-Agents—43%
OSWorld 2.0—9.3%
TheAgentCompany24%—
Cybench17.5%—
DeepResearch Bench—54.9%
OSWorld—72.1%
BALROG32.6%—
ExploitBench—23.6%
GBAEval—48.8%
GDP.pdf—18%
LMArena Search—1221
METR Time Horizons45.2%—
Vending-Bench 2—7,204

Reasoning Claude Sonnet 4.6 leads

Claude 3.5 Sonnet: 23.1 (#183), Claude Sonnet 4.6: 46.1 (#45)

Reasoning benchmarks
BenchmarkClaude 3.5 SonnetClaude Sonnet 4.6
LMArena Hard Prompts13051484
DTBench67.8%89.9%
Epoch Capabilities Index133.55152.24
ForecastBench60.762
ARC-AGI-2—60.4%
SimpleBench41.4%—
NYT Connections (extended)—80.9%
ARC-AGI-1—86.5%
CritPt—3.1%
Chess Puzzles—13%
EnigmaEval0.9%—
Thematic Generalization—76.3%
LiveBench Reasoning56.7%—
Mystery Game Puzzles—16%
LiveBench Data Analysis55%—
LMCA—46.5%
LiveBench59%—

Math Claude Sonnet 4.6 leads

Claude 3.5 Sonnet: 19.2 (#288), Claude Sonnet 4.6: 52.9 (#49)

Math benchmarks
BenchmarkClaude 3.5 SonnetClaude Sonnet 4.6
OTIS Mock AIME 2024-20258.5%85.8%
LMArena Math13071462
FrontierMath (Feb 2025 set)2.1%32.4%
FrontierMath Tier 4 (v1)0%8.3%
ProofBench—45%
Omni-MATH27.6%—
LiveBench Math52.3%—
MATH Level 556.9%—

Knowledge Claude Sonnet 4.6 leads

Claude 3.5 Sonnet: 28.6 (#245), Claude Sonnet 4.6: 51.7 (#65)

Knowledge benchmarks
BenchmarkClaude 3.5 SonnetClaude Sonnet 4.6
GPQA Diamond55.3%87.4%
LMArena Expert12651500
Humanity's Last Exam4.1%—
SimpleQA Verified—35.5%
MMLU-Pro77.7%—
Confabulations19.9%—
Vectara Hallucination Rate—10.6%
GPQA (HELM)56.5%—
MMLU87.3%—

Multimodal Claude Sonnet 4.6 leads

Claude 3.5 Sonnet: 26.5 (#120), Claude Sonnet 4.6: 38.0 (#68)

Multimodal benchmarks
BenchmarkClaude 3.5 SonnetClaude Sonnet 4.6
LMArena Vision11251283
Video-MME60%—
GeoBench62%—
VPCT33%—
Blueprint-Bench 2—6.7%
LMArena Document—1482

Multilingual Claude Sonnet 4.6 leads

Claude 3.5 Sonnet: 43.2 (#185), Claude Sonnet 4.6: 54.4 (#41)

Multilingual benchmarks
BenchmarkClaude 3.5 SonnetClaude Sonnet 4.6
LMArena Non-English12831440
LMArena Chinese12721491
LMArena French13051465
LMArena German12971428
LMArena Japanese12341420
LMArena Korean12001411
LMArena Russian13061440
LMArena Spanish12901464

Instruction Following Claude Sonnet 4.6 leads

Claude 3.5 Sonnet: 68.8 (#182), Claude Sonnet 4.6: 77.4 (#25)

Instruction Following benchmarks
BenchmarkClaude 3.5 SonnetClaude Sonnet 4.6
LMArena Instruction Following12971475
LiveBench Instruction Following69.3%—
IFEval85.5%—

Long Context Claude Sonnet 4.6 leads

Claude 3.5 Sonnet: 39.9 (#167), Claude Sonnet 4.6: 45.3 (#44)

Long Context benchmarks
BenchmarkClaude 3.5 SonnetClaude Sonnet 4.6
LMArena Longer Query13111479

Writing & Preference Claude Sonnet 4.6 leads

Claude 3.5 Sonnet: 52.9 (#164), Claude Sonnet 4.6: 70.2 (#22)

Writing & Preference benchmarks
BenchmarkClaude 3.5 SonnetClaude Sonnet 4.6
LMArena Text12981458
LMArena Creative Writing12921435
EQ-Bench Creative Writing14511810
LMArena Multi-Turn13261464
Short-Story Creative Writing80.3%—
WildBench79.2%—
EQ-Bench 4—1207
LiveBench Language53.8%—

Frequently asked questions

Is Claude 3.5 Sonnet better than Claude Sonnet 4.6?

Claude Sonnet 4.6 is the stronger model overall, scoring 50.3 to 34.6 on the Noometry Index.

Is Claude 3.5 Sonnet or Claude Sonnet 4.6 better for coding?

Claude Sonnet 4.6 scores higher on coding benchmarks: 46.3 versus 39.0 in the Noometry coding category.

How many benchmarks do Claude 3.5 Sonnet and Claude Sonnet 4.6 share?

27 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and Claude Sonnet 4.6 has 57.

Related comparisons

Go deeper