Model comparison

Claude 2 vs Claude Sonnet 4.6

Claude Sonnet 4.6 is the stronger model overall, scoring 50.3 to 25.0 on the Noometry Index.

Last verified . 4 shared benchmarks.

Claude 2 Anthropic

25.0

Rank #346 Reported

Claude Sonnet 4.6 Anthropic

50.3

Rank #50 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Claude 2 scores higher in 0 categories and Claude Sonnet 4.6 in 3 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Sonnet 4.6 leads 52.9 to 9.3.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 2.5% for Claude 2 and 85.8% for Claude Sonnet 4.6.

Side by side

Claude 2 and Claude Sonnet 4.6 specifications
Claude 2Claude Sonnet 4.6
ProviderAnthropicAnthropic
Noometry Index25.050.3
Released2023-07-112026-02-17
WeightsProprietaryProprietary
Context window—1M
Max output—128K
Input $ / M tokens—$3
Output $ / M tokens—$15
Results tracked857

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Claude 2: —, Claude Sonnet 4.6: 46.3 (#67)

Coding benchmarks
BenchmarkClaude 2Claude Sonnet 4.6
SWE-bench Verified—75.2%
DeepSWE—29.9%
FrontierCode—24.3%
LMArena WebDev—1522
SciCode—46.8%
WeirdML—66.1%
LMArena Coding—1504
ALE-Bench—1,327
HumanEval+61.6%—

Agentic & Tool Use Not comparable

Claude 2: —, Claude Sonnet 4.6: 39.1 (#28)

Agentic & Tool Use benchmarks
BenchmarkClaude 2Claude Sonnet 4.6
Terminal-Bench—53.4%
APEX-Agents—43%
OSWorld 2.0—9.3%
DeepResearch Bench—54.9%
OSWorld—72.1%
ExploitBench—23.6%
GBAEval—48.8%
GDP.pdf—18%
LMArena Search—1221
Vending-Bench 2—7,204

Reasoning Claude Sonnet 4.6 leads

Claude 2: 21.7 (#216), Claude Sonnet 4.6: 46.1 (#45)

Reasoning benchmarks
BenchmarkClaude 2Claude Sonnet 4.6
DTBench51.9%89.9%
Epoch Capabilities Index120.13152.24
ARC-AGI-2—60.4%
NYT Connections (extended)—80.9%
ARC-AGI-1—86.5%
CritPt—3.1%
Chess Puzzles—13%
Thematic Generalization—76.3%
LMArena Hard Prompts—1484
Mystery Game Puzzles—16%
LMCA—46.5%
ForecastBench—62

Math Claude Sonnet 4.6 leads

Claude 2: 9.3 (#320), Claude Sonnet 4.6: 52.9 (#49)

Math benchmarks
BenchmarkClaude 2Claude Sonnet 4.6
OTIS Mock AIME 2024-20252.5%85.8%
ProofBench—45%
LMArena Math—1462
MATH Level 511.7%—
FrontierMath (Feb 2025 set)—32.4%
FrontierMath Tier 4 (v1)—8.3%

Knowledge Claude Sonnet 4.6 leads

Claude 2: 16.9 (#287), Claude Sonnet 4.6: 51.7 (#65)

Knowledge benchmarks
BenchmarkClaude 2Claude Sonnet 4.6
GPQA Diamond34.7%87.4%
SimpleQA Verified—35.5%
Vectara Hallucination Rate—10.6%
LMArena Expert—1500
MMLU78.5%—
TriviaQA87.5%—

Multimodal Not comparable

Claude 2: —, Claude Sonnet 4.6: 38.0 (#68)

Multimodal benchmarks
BenchmarkClaude 2Claude Sonnet 4.6
LMArena Vision—1283
Blueprint-Bench 2—6.7%
LMArena Document—1482

Multilingual Not comparable

Claude 2: —, Claude Sonnet 4.6: 54.4 (#41)

Multilingual benchmarks
BenchmarkClaude 2Claude Sonnet 4.6
LMArena Non-English—1440
LMArena Chinese—1491
LMArena French—1465
LMArena German—1428
LMArena Japanese—1420
LMArena Korean—1411
LMArena Russian—1440
LMArena Spanish—1464

Instruction Following Not comparable

Claude 2: —, Claude Sonnet 4.6: 77.4 (#25)

Instruction Following benchmarks
BenchmarkClaude 2Claude Sonnet 4.6
LMArena Instruction Following—1475

Long Context Not comparable

Claude 2: —, Claude Sonnet 4.6: 45.3 (#44)

Long Context benchmarks
BenchmarkClaude 2Claude Sonnet 4.6
LMArena Longer Query—1479

Writing & Preference Not comparable

Claude 2: —, Claude Sonnet 4.6: 70.2 (#22)

Writing & Preference benchmarks
BenchmarkClaude 2Claude Sonnet 4.6
LMArena Text—1458
LMArena Creative Writing—1435
EQ-Bench Creative Writing—1810
EQ-Bench 4—1207
LMArena Multi-Turn—1464

Frequently asked questions

Is Claude 2 better than Claude Sonnet 4.6?

Claude Sonnet 4.6 is the stronger model overall, scoring 50.3 to 25.0 on the Noometry Index.

How many benchmarks do Claude 2 and Claude Sonnet 4.6 share?

4 benchmarks have published results for both models. Claude 2 has 8 scored results on Noometry and Claude Sonnet 4.6 has 57.

Related comparisons

Go deeper