Model comparison

Claude 2 vs Claude Sonnet 4.5

Claude Sonnet 4.5 is the stronger model overall, scoring 44.1 to 25.0 on the Noometry Index.

Last verified . 5 shared benchmarks.

Claude 2 Anthropic

25.0

Rank #346 Reported

Claude Sonnet 4.5 Anthropic

44.1

Rank #81 Confirmed

Summary

  • They share 5 benchmarks with published results for both. Claude 2 scores higher in 0 categories and Claude Sonnet 4.5 in 3 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Claude Sonnet 4.5 leads 48.4 to 16.9.
  • The biggest single-benchmark swing is MATH Level 5: 11.7% for Claude 2 and 97.7% for Claude Sonnet 4.5.

Side by side

Claude 2 and Claude Sonnet 4.5 specifications
Claude 2Claude Sonnet 4.5
ProviderAnthropicAnthropic
Noometry Index25.044.1
Released2023-07-112025-09-29
WeightsProprietaryProprietary
Context window—200K
Max output—64K
Input $ / M tokens—$3
Output $ / M tokens—$15
Results tracked873

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Claude 2: —, Claude Sonnet 4.5: 47.3 (#61)

Coding benchmarks
BenchmarkClaude 2Claude Sonnet 4.5
SWE-bench Verified—71.3%
SWE-bench Verified (bash only)—71.4%
LMArena WebDev—1393
SWE-bench Multilingual—67%
SciCode—44.7%
GSO—14.7%
WeirdML—47.7%
LMArena Coding—1489
ALE-Bench—796.15
AlgoTune—1.52
HumanEval+61.6%—

Agentic & Tool Use Not comparable

Claude 2: —, Claude Sonnet 4.5: 38.3 (#32)

Agentic & Tool Use benchmarks
BenchmarkClaude 2Claude Sonnet 4.5
Terminal-Bench—46.5%
Berkeley Function Calling Leaderboard—73.2%
GDPval—42.5%
Remote Labor Index—2.1%
τ²-bench Airline—72%
τ²-bench Banking—25.3%
τ²-bench Retail—72.4%
τ²-bench Telecom—84.9%
Cybench—60%
DeepResearch Bench—52.6%
OSWorld—62.9%
LMArena Search—1159
METR Time Horizons—67.4%
Vending-Bench 2—3,839

Reasoning Claude Sonnet 4.5 leads

Claude 2: 21.7 (#216), Claude Sonnet 4.5: 26.9 (#125)

Reasoning benchmarks
BenchmarkClaude 2Claude Sonnet 4.5
DTBench51.9%83.2%
Epoch Capabilities Index120.13146.84
ARC-AGI-2—13.6%
SimpleBench—54.3%
Kagi LLM Benchmark—57.9%
NYT Connections (extended)—37.3%
ARC-AGI-1—63.7%
CritPt—1.1%
Chess Puzzles—12%
EnigmaEval—6%
EBR-Bench—2.4%
LMArena Hard Prompts—1462
Mystery Game Puzzles—17%
LMCA—38.8%
ForecastBench—61.9

Math Claude Sonnet 4.5 leads

Claude 2: 9.3 (#320), Claude Sonnet 4.5: 32.3 (#216)

Math benchmarks
BenchmarkClaude 2Claude Sonnet 4.5
OTIS Mock AIME 2024-20252.5%77.8%
MATH Level 511.7%97.7%
FrontierMath (Tiers 1-3)—23.9%
FrontierMath Tier 4—2.4%
ProofBench—19%
Omni-MATH—55.3%
LMArena Math—1449
FrontierMath (Feb 2025 set)—15.2%
FrontierMath Tier 4 (v1)—4.2%

Knowledge Claude Sonnet 4.5 leads

Claude 2: 16.9 (#287), Claude Sonnet 4.5: 48.4 (#76)

Knowledge benchmarks
BenchmarkClaude 2Claude Sonnet 4.5
GPQA Diamond34.7%82.3%
Humanity's Last Exam—13.7%
SimpleQA Verified—30.7%
MMLU-Pro—86.9%
Vectara Hallucination Rate—12%
GPQA (HELM)—68.6%
LMArena Expert—1482
MMLU78.5%—
TriviaQA87.5%—

Multimodal Not comparable

Claude 2: —, Claude Sonnet 4.5: 34.8 (#89)

Multimodal benchmarks
BenchmarkClaude 2Claude Sonnet 4.5
VPCT—39.8%
LMArena Document—1450

Multilingual Not comparable

Claude 2: —, Claude Sonnet 4.5: 53.4 (#69)

Multilingual benchmarks
BenchmarkClaude 2Claude Sonnet 4.5
LMArena Non-English—1425
LMArena Chinese—1459
LMArena French—1458
LMArena German—1427
LMArena Japanese—1390
LMArena Korean—1403
LMArena Russian—1437
LMArena Spanish—1457

Instruction Following Not comparable

Claude 2: —, Claude Sonnet 4.5: 75.0 (#78)

Instruction Following benchmarks
BenchmarkClaude 2Claude Sonnet 4.5
IFEval—85%
LMArena Instruction Following—1459

Long Context Not comparable

Claude 2: —, Claude Sonnet 4.5: 45.2 (#46)

Long Context benchmarks
BenchmarkClaude 2Claude Sonnet 4.5
LMArena Longer Query—1476

Writing & Preference Not comparable

Claude 2: —, Claude Sonnet 4.5: 66.5 (#34)

Writing & Preference benchmarks
BenchmarkClaude 2Claude Sonnet 4.5
LMArena Text—1439
LMArena Creative Writing—1442
EQ-Bench Creative Writing—1678
WildBench—85.4%
LMArena Multi-Turn—1465

Frequently asked questions

Is Claude 2 better than Claude Sonnet 4.5?

Claude Sonnet 4.5 is the stronger model overall, scoring 44.1 to 25.0 on the Noometry Index.

How many benchmarks do Claude 2 and Claude Sonnet 4.5 share?

5 benchmarks have published results for both models. Claude 2 has 8 scored results on Noometry and Claude Sonnet 4.5 has 73.

Related comparisons

Go deeper