Model comparison

Claude 2.1 vs Claude 3.5 Sonnet

Claude 3.5 Sonnet is the stronger model overall, scoring 34.6 to 25.2 on the Noometry Index.

Last verified . 7 shared benchmarks.

Claude 2.1 Anthropic

25.2

Rank #345 Reported

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

Summary

  • They share 7 benchmarks with published results for both. Claude 2.1 scores higher in 0 categories and Claude 3.5 Sonnet in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Claude 3.5 Sonnet leads 28.6 to 15.4.
  • The biggest single-benchmark swing is WeirdML: 7.1% for Claude 2.1 and 40% for Claude 3.5 Sonnet.

Side by side

Claude 2.1 and Claude 3.5 Sonnet specifications
Claude 2.1Claude 3.5 Sonnet
ProviderAnthropicAnthropic
Noometry Index25.234.6
Released2023-11-212024-06-20
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked760

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.5 Sonnet leads

Claude 2.1: 26.2 (#327), Claude 3.5 Sonnet: 39.0 (#165)

Coding benchmarks
BenchmarkClaude 2.1Claude 3.5 Sonnet
WeirdML7.1%40%
Aider Polyglot—51.6%
GSO—4.6%
BigCodeBench Instruct—46.8%
LiveBench Coding—67.1%
LMArena Coding—1342
BigCodeBench Complete—58.6%
CadEval—48%
HumanEval+—81.7%
MBPP+—74.3%

Agentic & Tool Use Not comparable

Claude 2.1: —, Claude 3.5 Sonnet: 32.3 (#67)

Agentic & Tool Use benchmarks
BenchmarkClaude 2.1Claude 3.5 Sonnet
TheAgentCompany—24%
Cybench—17.5%
BALROG—32.6%
METR Time Horizons—45.2%

Reasoning Claude 3.5 Sonnet leads

Claude 2.1: 21.4 (#221), Claude 3.5 Sonnet: 23.1 (#183)

Reasoning benchmarks
BenchmarkClaude 2.1Claude 3.5 Sonnet
DTBench51%67.8%
Epoch Capabilities Index119.27133.55
ForecastBench54.260.7
SimpleBench—41.4%
EnigmaEval—0.9%
LiveBench Reasoning—56.7%
LMArena Hard Prompts—1305
LiveBench Data Analysis—55%
LiveBench—59%

Math Claude 3.5 Sonnet leads

Claude 2.1: 10.2 (#315), Claude 3.5 Sonnet: 19.2 (#288)

Math benchmarks
BenchmarkClaude 2.1Claude 3.5 Sonnet
OTIS Mock AIME 2024-20251.9%8.5%
Omni-MATH—27.6%
LiveBench Math—52.3%
LMArena Math—1307
MATH Level 5—56.9%
FrontierMath (Feb 2025 set)—2.1%
FrontierMath Tier 4 (v1)—0%

Knowledge Claude 3.5 Sonnet leads

Claude 2.1: 15.4 (#292), Claude 3.5 Sonnet: 28.6 (#245)

Knowledge benchmarks
BenchmarkClaude 2.1Claude 3.5 Sonnet
GPQA Diamond33%55.3%
MMLU73.5%87.3%
Humanity's Last Exam—4.1%
MMLU-Pro—77.7%
Confabulations—19.9%
GPQA (HELM)—56.5%
LMArena Expert—1265

Multimodal Not comparable

Claude 2.1: —, Claude 3.5 Sonnet: 26.5 (#120)

Multimodal benchmarks
BenchmarkClaude 2.1Claude 3.5 Sonnet
LMArena Vision—1125
Video-MME—60%
GeoBench—62%
VPCT—33%

Multilingual Not comparable

Claude 2.1: —, Claude 3.5 Sonnet: 43.2 (#185)

Multilingual benchmarks
BenchmarkClaude 2.1Claude 3.5 Sonnet
LMArena Non-English—1283
LMArena Chinese—1272
LMArena French—1305
LMArena German—1297
LMArena Japanese—1234
LMArena Korean—1200
LMArena Russian—1306
LMArena Spanish—1290

Instruction Following Not comparable

Claude 2.1: —, Claude 3.5 Sonnet: 68.8 (#182)

Instruction Following benchmarks
BenchmarkClaude 2.1Claude 3.5 Sonnet
LiveBench Instruction Following—69.3%
IFEval—85.5%
LMArena Instruction Following—1297

Long Context Not comparable

Claude 2.1: —, Claude 3.5 Sonnet: 39.9 (#167)

Long Context benchmarks
BenchmarkClaude 2.1Claude 3.5 Sonnet
LMArena Longer Query—1311

Writing & Preference Not comparable

Claude 2.1: —, Claude 3.5 Sonnet: 52.9 (#164)

Writing & Preference benchmarks
BenchmarkClaude 2.1Claude 3.5 Sonnet
LMArena Text—1298
LMArena Creative Writing—1292
Short-Story Creative Writing—80.3%
EQ-Bench Creative Writing—1451
WildBench—79.2%
LMArena Multi-Turn—1326
LiveBench Language—53.8%

Frequently asked questions

Is Claude 2.1 better than Claude 3.5 Sonnet?

Claude 3.5 Sonnet is the stronger model overall, scoring 34.6 to 25.2 on the Noometry Index.

Is Claude 2.1 or Claude 3.5 Sonnet better for coding?

Claude 3.5 Sonnet scores higher on coding benchmarks: 39.0 versus 26.2 in the Noometry coding category.

How many benchmarks do Claude 2.1 and Claude 3.5 Sonnet share?

7 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and Claude 3.5 Sonnet has 60.

Related comparisons

Go deeper