Model comparison

Claude 3 Sonnet vs Claude Sonnet 4

Claude Sonnet 4 is the stronger model overall, scoring 40.8 to 29.0 on the Noometry Index.

Last verified . 24 shared benchmarks.

Claude 3 Sonnet Anthropic

29.0

Rank #319 Confirmed

Claude Sonnet 4 Anthropic

40.8

Rank #145 Confirmed

Summary

  • They share 24 benchmarks with published results for both. Claude 3 Sonnet scores higher in 1 category and Claude Sonnet 4 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Sonnet 4 leads 43.3 to 10.7.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 2.5% for Claude 3 Sonnet and 71.1% for Claude Sonnet 4.

Side by side

Claude 3 Sonnet and Claude Sonnet 4 specifications
Claude 3 SonnetClaude Sonnet 4
ProviderAnthropicAnthropic
Noometry Index29.040.8
Released2024-02-292025-05-22
WeightsProprietaryProprietary
Context window—200K
Max output—64K
Input $ / M tokens—$3
Output $ / M tokens—$15
Results tracked3058

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Sonnet 4 leads

Claude 3 Sonnet: 29.6 (#302), Claude Sonnet 4: 43.5 (#88)

Coding benchmarks
BenchmarkClaude 3 SonnetClaude Sonnet 4
WeirdML10.2%46.1%
LMArena Coding12231414
SWE-bench Verified (bash only)—64.9%
Aider Polyglot—61.3%
SciCode—40%
GSO—4.9%
BigCodeBench Instruct42.7%—
BigCodeBench Complete53.8%—
ALE-Bench—655.35
HumanEval+64%—
MBPP+69.3%—

Agentic & Tool Use Not comparable

Claude 3 Sonnet: —, Claude Sonnet 4: 38.5 (#31)

Agentic & Tool Use benchmarks
BenchmarkClaude 3 SonnetClaude Sonnet 4
TheAgentCompany—33.1%
Cybench—35%
DeepResearch Bench—46.6%
OSWorld—43.9%
METR Time Horizons—62%

Reasoning Claude Sonnet 4 leads

Claude 3 Sonnet: 20.5 (#237), Claude Sonnet 4: 22.9 (#187)

Reasoning benchmarks
BenchmarkClaude 3 SonnetClaude Sonnet 4
LMArena Hard Prompts11971372
DTBench53.6%77.1%
Epoch Capabilities Index120.7141.69
ARC-AGI-2—5.9%
SimpleBench—45.5%
Kagi LLM Benchmark—73%
ARC-AGI-1—40%
CritPt—0.3%
EnigmaEval—3.1%
LMCA—29%
ForecastBench—60.2
WinoGrande75.1%—

Math Claude Sonnet 4 leads

Claude 3 Sonnet: 10.7 (#310), Claude Sonnet 4: 43.3 (#80)

Math benchmarks
BenchmarkClaude 3 SonnetClaude Sonnet 4
OTIS Mock AIME 2024-20252.5%71.1%
LMArena Math12131375
MATH Level 518.2%84.4%
Omni-MATH—60.2%
FrontierMath (Feb 2025 set)—4.1%
FrontierMath Tier 4 (v1)—0%

Knowledge Claude Sonnet 4 leads

Claude 3 Sonnet: 21.1 (#276), Claude Sonnet 4: 41.8 (#108)

Knowledge benchmarks
BenchmarkClaude 3 SonnetClaude Sonnet 4
GPQA Diamond40.6%79.2%
LMArena Expert11731372
Humanity's Last Exam—7.8%
MMLU-Pro—84.3%
Confabulations—13.2%
Vectara Hallucination Rate—10.3%
GPQA (HELM)—70.6%
MMLU75.9%—

Multimodal Too close to call

Claude 3 Sonnet: 25.2 (#125), Claude Sonnet 4: 26.2 (#121)

Multimodal benchmarks
BenchmarkClaude 3 SonnetClaude Sonnet 4
LMArena Vision9841191
GeoBench—37%
VPCT—34%
MindCube—44.8%

Multilingual Claude Sonnet 4 leads

Claude 3 Sonnet: 37.8 (#234), Claude Sonnet 4: 46.7 (#156)

Multilingual benchmarks
BenchmarkClaude 3 SonnetClaude Sonnet 4
LMArena Non-English12051333
LMArena Chinese11891350
LMArena French12291363
LMArena German12041331
LMArena Japanese11311302
LMArena Korean11281291
LMArena Russian12271355
LMArena Spanish12041357

Instruction Following Claude Sonnet 4 leads

Claude 3 Sonnet: 62.8 (#235), Claude Sonnet 4: 71.7 (#145)

Instruction Following benchmarks
BenchmarkClaude 3 SonnetClaude Sonnet 4
LMArena Instruction Following11991376
IFEval—84%

Long Context Claude 3 Sonnet leads

Claude 3 Sonnet: 36.7 (#228), Claude Sonnet 4: 33.7 (#259)

Long Context benchmarks
BenchmarkClaude 3 SonnetClaude Sonnet 4
LMArena Longer Query12111398
Fiction.LiveBench—46.9%

Writing & Preference Claude Sonnet 4 leads

Claude 3 Sonnet: 42.1 (#238), Claude Sonnet 4: 57.1 (#132)

Writing & Preference benchmarks
BenchmarkClaude 3 SonnetClaude Sonnet 4
LMArena Text12181351
LMArena Creative Writing11861345
LMArena Multi-Turn12271376
Short-Story Creative Writing—81.4%
EQ-Bench Creative Writing—1483
WildBench—83.8%

Frequently asked questions

Is Claude 3 Sonnet better than Claude Sonnet 4?

Claude Sonnet 4 is the stronger model overall, scoring 40.8 to 29.0 on the Noometry Index.

Is Claude 3 Sonnet or Claude Sonnet 4 better for coding?

Claude Sonnet 4 scores higher on coding benchmarks: 43.5 versus 29.6 in the Noometry coding category.

How many benchmarks do Claude 3 Sonnet and Claude Sonnet 4 share?

24 benchmarks have published results for both models. Claude 3 Sonnet has 30 scored results on Noometry and Claude Sonnet 4 has 58.

Related comparisons

Go deeper