Model comparison

Claude 2.1 vs Claude Sonnet 4

Claude Sonnet 4 is the stronger model overall, scoring 40.8 to 25.2 on the Noometry Index.

Last verified . 6 shared benchmarks.

Claude 2.1 Anthropic

25.2

Rank #345 Reported

Claude Sonnet 4 Anthropic

40.8

Rank #145 Confirmed

Summary

  • They share 6 benchmarks with published results for both. Claude 2.1 scores higher in 0 categories and Claude Sonnet 4 in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Sonnet 4 leads 43.3 to 10.2.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 1.9% for Claude 2.1 and 71.1% for Claude Sonnet 4.

Side by side

Claude 2.1 and Claude Sonnet 4 specifications
Claude 2.1Claude Sonnet 4
ProviderAnthropicAnthropic
Noometry Index25.240.8
Released2023-11-212025-05-22
WeightsProprietaryProprietary
Context window—200K
Max output—64K
Input $ / M tokens—$3
Output $ / M tokens—$15
Results tracked758

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Sonnet 4 leads

Claude 2.1: 26.2 (#327), Claude Sonnet 4: 43.5 (#88)

Coding benchmarks
BenchmarkClaude 2.1Claude Sonnet 4
WeirdML7.1%46.1%
SWE-bench Verified (bash only)—64.9%
Aider Polyglot—61.3%
SciCode—40%
GSO—4.9%
LMArena Coding—1414
ALE-Bench—655.35

Agentic & Tool Use Not comparable

Claude 2.1: —, Claude Sonnet 4: 38.5 (#31)

Agentic & Tool Use benchmarks
BenchmarkClaude 2.1Claude Sonnet 4
TheAgentCompany—33.1%
Cybench—35%
DeepResearch Bench—46.6%
OSWorld—43.9%
METR Time Horizons—62%

Reasoning Claude Sonnet 4 leads

Claude 2.1: 21.4 (#221), Claude Sonnet 4: 22.9 (#187)

Reasoning benchmarks
BenchmarkClaude 2.1Claude Sonnet 4
DTBench51%77.1%
Epoch Capabilities Index119.27141.69
ForecastBench54.260.2
ARC-AGI-2—5.9%
SimpleBench—45.5%
Kagi LLM Benchmark—73%
ARC-AGI-1—40%
CritPt—0.3%
EnigmaEval—3.1%
LMArena Hard Prompts—1372
LMCA—29%

Math Claude Sonnet 4 leads

Claude 2.1: 10.2 (#315), Claude Sonnet 4: 43.3 (#80)

Math benchmarks
BenchmarkClaude 2.1Claude Sonnet 4
OTIS Mock AIME 2024-20251.9%71.1%
Omni-MATH—60.2%
LMArena Math—1375
MATH Level 5—84.4%
FrontierMath (Feb 2025 set)—4.1%
FrontierMath Tier 4 (v1)—0%

Knowledge Claude Sonnet 4 leads

Claude 2.1: 15.4 (#292), Claude Sonnet 4: 41.8 (#108)

Knowledge benchmarks
BenchmarkClaude 2.1Claude Sonnet 4
GPQA Diamond33%79.2%
Humanity's Last Exam—7.8%
MMLU-Pro—84.3%
Confabulations—13.2%
Vectara Hallucination Rate—10.3%
GPQA (HELM)—70.6%
LMArena Expert—1372
MMLU73.5%—

Multimodal Not comparable

Claude 2.1: —, Claude Sonnet 4: 26.2 (#121)

Multimodal benchmarks
BenchmarkClaude 2.1Claude Sonnet 4
LMArena Vision—1191
GeoBench—37%
VPCT—34%
MindCube—44.8%

Multilingual Not comparable

Claude 2.1: —, Claude Sonnet 4: 46.7 (#156)

Multilingual benchmarks
BenchmarkClaude 2.1Claude Sonnet 4
LMArena Non-English—1333
LMArena Chinese—1350
LMArena French—1363
LMArena German—1331
LMArena Japanese—1302
LMArena Korean—1291
LMArena Russian—1355
LMArena Spanish—1357

Instruction Following Not comparable

Claude 2.1: —, Claude Sonnet 4: 71.7 (#145)

Instruction Following benchmarks
BenchmarkClaude 2.1Claude Sonnet 4
IFEval—84%
LMArena Instruction Following—1376

Long Context Not comparable

Claude 2.1: —, Claude Sonnet 4: 33.7 (#259)

Long Context benchmarks
BenchmarkClaude 2.1Claude Sonnet 4
Fiction.LiveBench—46.9%
LMArena Longer Query—1398

Writing & Preference Not comparable

Claude 2.1: —, Claude Sonnet 4: 57.1 (#132)

Writing & Preference benchmarks
BenchmarkClaude 2.1Claude Sonnet 4
LMArena Text—1351
LMArena Creative Writing—1345
Short-Story Creative Writing—81.4%
EQ-Bench Creative Writing—1483
WildBench—83.8%
LMArena Multi-Turn—1376

Frequently asked questions

Is Claude 2.1 better than Claude Sonnet 4?

Claude Sonnet 4 is the stronger model overall, scoring 40.8 to 25.2 on the Noometry Index.

Is Claude 2.1 or Claude Sonnet 4 better for coding?

Claude Sonnet 4 scores higher on coding benchmarks: 43.5 versus 26.2 in the Noometry coding category.

How many benchmarks do Claude 2.1 and Claude Sonnet 4 share?

6 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and Claude Sonnet 4 has 58.

Related comparisons

Go deeper