Model comparison

Claude 2 vs Claude 3.7 Sonnet

Claude 3.7 Sonnet is the stronger model overall, scoring 39.5 to 25.0 on the Noometry Index.

Last verified . 4 shared benchmarks.

Claude 2 Anthropic

25.0

Rank #346 Reported

Claude 3.7 Sonnet Anthropic

39.5

Rank #164 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Claude 2 scores higher in 1 category and Claude 3.7 Sonnet in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude 3.7 Sonnet leads 37.5 to 9.3.
  • The biggest single-benchmark swing is MATH Level 5: 11.7% for Claude 2 and 91.2% for Claude 3.7 Sonnet.

Side by side

Claude 2 and Claude 3.7 Sonnet specifications
Claude 2Claude 3.7 Sonnet
ProviderAnthropicAnthropic
Noometry Index25.039.5
Released2023-07-112025-02-24
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked858

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Claude 2: —, Claude 3.7 Sonnet: 40.6 (#136)

Coding benchmarks
BenchmarkClaude 2Claude 3.7 Sonnet
SWE-bench Verified—61%
SWE-bench Verified (bash only)—52.8%
Aider Polyglot—64.9%
GSO—3.8%
LiveBench Coding—74.5%
LMArena Coding—1361
CadEval—54%
HumanEval+61.6%—

Agentic & Tool Use Not comparable

Claude 2: —, Claude 3.7 Sonnet: 34.1 (#50)

Agentic & Tool Use benchmarks
BenchmarkClaude 2Claude 3.7 Sonnet
TheAgentCompany—30.9%
Cybench—20%
DeepResearch Bench—43.6%
OSWorld—35.8%
METR Time Horizons—60%

Reasoning Claude 2 leads

Claude 2: 21.7 (#216), Claude 3.7 Sonnet: 18.6 (#277)

Reasoning benchmarks
BenchmarkClaude 2Claude 3.7 Sonnet
Epoch Capabilities Index120.13141.16
ARC-AGI-2—0.9%
SimpleBench—46.4%
ARC-AGI-1—28.6%
EnigmaEval—4.2%
LiveBench Reasoning—87.8%
LMArena Hard Prompts—1333
DTBench51.9%—
LiveBench Data Analysis—74%
ForecastBench—61.8
LiveBench—76.1%

Math Claude 3.7 Sonnet leads

Claude 2: 9.3 (#320), Claude 3.7 Sonnet: 37.5 (#153)

Math benchmarks
BenchmarkClaude 2Claude 3.7 Sonnet
OTIS Mock AIME 2024-20252.5%57.8%
MATH Level 511.7%91.2%
Omni-MATH—33%
LiveBench Math—79%
LMArena Math—1337
FrontierMath (Feb 2025 set)—4.1%

Knowledge Claude 3.7 Sonnet leads

Claude 2: 16.9 (#287), Claude 3.7 Sonnet: 39.8 (#130)

Knowledge benchmarks
BenchmarkClaude 2Claude 3.7 Sonnet
GPQA Diamond34.7%79.7%
Humanity's Last Exam—8%
MMLU-Pro—78.4%
Confabulations—14.7%
GPQA (HELM)—60.8%
LMArena Expert—1321
MMLU78.5%—
TriviaQA87.5%—

Multimodal Not comparable

Claude 2: —, Claude 3.7 Sonnet: 33.7 (#95)

Multimodal benchmarks
BenchmarkClaude 2Claude 3.7 Sonnet
LMArena Vision—1169
GeoBench—68%
VPCT—39%
SpatialViz-Bench—33.9%

Multilingual Not comparable

Claude 2: —, Claude 3.7 Sonnet: 44.1 (#179)

Multilingual benchmarks
BenchmarkClaude 2Claude 3.7 Sonnet
LMArena Non-English—1296
LMArena Chinese—1299
LMArena French—1303
LMArena German—1301
LMArena Japanese—1267
LMArena Korean—1249
LMArena Russian—1311
LMArena Spanish—1298

Instruction Following Not comparable

Claude 2: —, Claude 3.7 Sonnet: 72.9 (#125)

Instruction Following benchmarks
BenchmarkClaude 2Claude 3.7 Sonnet
LiveBench Instruction Following—81.3%
IFEval—83.4%
LMArena Instruction Following—1352

Long Context Not comparable

Claude 2: —, Claude 3.7 Sonnet: 50.3 (#10)

Long Context benchmarks
BenchmarkClaude 2Claude 3.7 Sonnet
Fiction.LiveBench—83.3%
LMArena Longer Query—1373

Writing & Preference Not comparable

Claude 2: —, Claude 3.7 Sonnet: 54.4 (#150)

Writing & Preference benchmarks
BenchmarkClaude 2Claude 3.7 Sonnet
LMArena Text—1314
LMArena Creative Writing—1332
Short-Story Creative Writing—81.1%
EQ-Bench Creative Writing—1412
WildBench—81.4%
LMArena Multi-Turn—1339
LiveBench Language—59.9%

Frequently asked questions

Is Claude 2 better than Claude 3.7 Sonnet?

Claude 3.7 Sonnet is the stronger model overall, scoring 39.5 to 25.0 on the Noometry Index.

How many benchmarks do Claude 2 and Claude 3.7 Sonnet share?

4 benchmarks have published results for both models. Claude 2 has 8 scored results on Noometry and Claude 3.7 Sonnet has 58.

Related comparisons

Go deeper