Model comparison

Claude 2 vs Claude Sonnet 4

Claude Sonnet 4 is the stronger model overall, scoring 40.8 to 25.0 on the Noometry Index.

Last verified . 5 shared benchmarks.

Claude 2 Anthropic

25.0

Rank #346 Reported

Claude Sonnet 4 Anthropic

40.8

Rank #145 Confirmed

Summary

  • They share 5 benchmarks with published results for both. Claude 2 scores higher in 0 categories and Claude Sonnet 4 in 3 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Sonnet 4 leads 43.3 to 9.3.
  • The biggest single-benchmark swing is MATH Level 5: 11.7% for Claude 2 and 84.4% for Claude Sonnet 4.

Side by side

Claude 2 and Claude Sonnet 4 specifications
Claude 2Claude Sonnet 4
ProviderAnthropicAnthropic
Noometry Index25.040.8
Released2023-07-112025-05-22
WeightsProprietaryProprietary
Context window—200K
Max output—64K
Input $ / M tokens—$3
Output $ / M tokens—$15
Results tracked858

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Claude 2: —, Claude Sonnet 4: 43.5 (#88)

Coding benchmarks
BenchmarkClaude 2Claude Sonnet 4
SWE-bench Verified (bash only)—64.9%
Aider Polyglot—61.3%
SciCode—40%
GSO—4.9%
WeirdML—46.1%
LMArena Coding—1414
ALE-Bench—655.35
HumanEval+61.6%—

Agentic & Tool Use Not comparable

Claude 2: —, Claude Sonnet 4: 38.5 (#31)

Agentic & Tool Use benchmarks
BenchmarkClaude 2Claude Sonnet 4
TheAgentCompany—33.1%
Cybench—35%
DeepResearch Bench—46.6%
OSWorld—43.9%
METR Time Horizons—62%

Reasoning Claude Sonnet 4 leads

Claude 2: 21.7 (#216), Claude Sonnet 4: 22.9 (#187)

Reasoning benchmarks
BenchmarkClaude 2Claude Sonnet 4
DTBench51.9%77.1%
Epoch Capabilities Index120.13141.69
ARC-AGI-2—5.9%
SimpleBench—45.5%
Kagi LLM Benchmark—73%
ARC-AGI-1—40%
CritPt—0.3%
EnigmaEval—3.1%
LMArena Hard Prompts—1372
LMCA—29%
ForecastBench—60.2

Math Claude Sonnet 4 leads

Claude 2: 9.3 (#320), Claude Sonnet 4: 43.3 (#80)

Math benchmarks
BenchmarkClaude 2Claude Sonnet 4
OTIS Mock AIME 2024-20252.5%71.1%
MATH Level 511.7%84.4%
Omni-MATH—60.2%
LMArena Math—1375
FrontierMath (Feb 2025 set)—4.1%
FrontierMath Tier 4 (v1)—0%

Knowledge Claude Sonnet 4 leads

Claude 2: 16.9 (#287), Claude Sonnet 4: 41.8 (#108)

Knowledge benchmarks
BenchmarkClaude 2Claude Sonnet 4
GPQA Diamond34.7%79.2%
Humanity's Last Exam—7.8%
MMLU-Pro—84.3%
Confabulations—13.2%
Vectara Hallucination Rate—10.3%
GPQA (HELM)—70.6%
LMArena Expert—1372
MMLU78.5%—
TriviaQA87.5%—

Multimodal Not comparable

Claude 2: —, Claude Sonnet 4: 26.2 (#121)

Multimodal benchmarks
BenchmarkClaude 2Claude Sonnet 4
LMArena Vision—1191
GeoBench—37%
VPCT—34%
MindCube—44.8%

Multilingual Not comparable

Claude 2: —, Claude Sonnet 4: 46.7 (#156)

Multilingual benchmarks
BenchmarkClaude 2Claude Sonnet 4
LMArena Non-English—1333
LMArena Chinese—1350
LMArena French—1363
LMArena German—1331
LMArena Japanese—1302
LMArena Korean—1291
LMArena Russian—1355
LMArena Spanish—1357

Instruction Following Not comparable

Claude 2: —, Claude Sonnet 4: 71.7 (#145)

Instruction Following benchmarks
BenchmarkClaude 2Claude Sonnet 4
IFEval—84%
LMArena Instruction Following—1376

Long Context Not comparable

Claude 2: —, Claude Sonnet 4: 33.7 (#259)

Long Context benchmarks
BenchmarkClaude 2Claude Sonnet 4
Fiction.LiveBench—46.9%
LMArena Longer Query—1398

Writing & Preference Not comparable

Claude 2: —, Claude Sonnet 4: 57.1 (#132)

Writing & Preference benchmarks
BenchmarkClaude 2Claude Sonnet 4
LMArena Text—1351
LMArena Creative Writing—1345
Short-Story Creative Writing—81.4%
EQ-Bench Creative Writing—1483
WildBench—83.8%
LMArena Multi-Turn—1376

Frequently asked questions

Is Claude 2 better than Claude Sonnet 4?

Claude Sonnet 4 is the stronger model overall, scoring 40.8 to 25.0 on the Noometry Index.

How many benchmarks do Claude 2 and Claude Sonnet 4 share?

5 benchmarks have published results for both models. Claude 2 has 8 scored results on Noometry and Claude Sonnet 4 has 58.

Related comparisons

Go deeper