Model comparison

Claude 2 vs Claude Haiku 4.5

Claude Haiku 4.5 is the stronger model overall, scoring 39.5 to 25.0 on the Noometry Index.

Last verified . 5 shared benchmarks.

Claude 2 Anthropic

25.0

Rank #346 Reported

Claude Haiku 4.5 Anthropic

39.5

Rank #165 Confirmed

Summary

  • They share 5 benchmarks with published results for both. Claude 2 scores higher in 1 category and Claude Haiku 4.5 in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Haiku 4.5 leads 44.9 to 9.3.
  • The biggest single-benchmark swing is MATH Level 5: 11.7% for Claude 2 and 96.4% for Claude Haiku 4.5.

Side by side

Claude 2 and Claude Haiku 4.5 specifications
Claude 2Claude Haiku 4.5
ProviderAnthropicAnthropic
Noometry Index25.039.5
Released2023-07-112025-10-15
WeightsProprietaryProprietary
Context window—200K
Max output—64K
Input $ / M tokens—$1
Output $ / M tokens—$5
Results tracked853

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Claude 2: —, Claude Haiku 4.5: 44.0 (#78)

Coding benchmarks
BenchmarkClaude 2Claude Haiku 4.5
SWE-bench Verified (bash only)—66.6%
LMArena WebDev—1330
SWE-bench Multilingual—64.7%
SciCode—43.3%
WeirdML—45.4%
LMArena Coding—1453
ALE-Bench—653.48
HumanEval+61.6%—

Agentic & Tool Use Not comparable

Claude 2: —, Claude Haiku 4.5: 33.6 (#52)

Agentic & Tool Use benchmarks
BenchmarkClaude 2Claude Haiku 4.5
Terminal-Bench—35.5%
Berkeley Function Calling Leaderboard—68.7%
DeepResearch Bench—45.5%
BALROG—31.2%
ExploitBench—13.7%
Vending-Bench 2—458.89

Reasoning Claude 2 leads

Claude 2: 21.7 (#216), Claude Haiku 4.5: 15.1 (#320)

Reasoning benchmarks
BenchmarkClaude 2Claude Haiku 4.5
DTBench51.9%73.6%
Epoch Capabilities Index120.13142.41
ARC-AGI-2—4%
NYT Connections (extended)—14.3%
ARC-AGI-1—47.7%
CritPt—0%
Chess Puzzles—8%
LMArena Hard Prompts—1420
LMCA—30.9%
ForecastBench—61.4

Math Claude Haiku 4.5 leads

Claude 2: 9.3 (#320), Claude Haiku 4.5: 44.9 (#78)

Math benchmarks
BenchmarkClaude 2Claude Haiku 4.5
OTIS Mock AIME 2024-20252.5%66.7%
MATH Level 511.7%96.4%
Omni-MATH—56.1%
LMArena Math—1396
FrontierMath (Feb 2025 set)—5.9%
FrontierMath Tier 4 (v1)—2.1%

Knowledge Claude Haiku 4.5 leads

Claude 2: 16.9 (#287), Claude Haiku 4.5: 37.7 (#153)

Knowledge benchmarks
BenchmarkClaude 2Claude Haiku 4.5
GPQA Diamond34.7%71.2%
SimpleQA Verified—13.2%
MMLU-Pro—77.7%
Vectara Hallucination Rate—9.8%
GPQA (HELM)—60.5%
LMArena Expert—1442
MMLU78.5%—
TriviaQA87.5%—

Multimodal Not comparable

Claude 2: —, Claude Haiku 4.5: 26.8 (#118)

Multimodal benchmarks
BenchmarkClaude 2Claude Haiku 4.5
Blueprint-Bench 2—0%
LMArena Document—1420

Multilingual Not comparable

Claude 2: —, Claude Haiku 4.5: 49.9 (#129)

Multilingual benchmarks
BenchmarkClaude 2Claude Haiku 4.5
LMArena Non-English—1377
LMArena Chinese—1417
LMArena French—1408
LMArena German—1375
LMArena Japanese—1339
LMArena Korean—1347
LMArena Russian—1381
LMArena Spanish—1420

Instruction Following Not comparable

Claude 2: —, Claude Haiku 4.5: 71.4 (#149)

Instruction Following benchmarks
BenchmarkClaude 2Claude Haiku 4.5
IFEval—80.1%
LMArena Instruction Following—1414

Long Context Not comparable

Claude 2: —, Claude Haiku 4.5: 43.6 (#92)

Long Context benchmarks
BenchmarkClaude 2Claude Haiku 4.5
LMArena Longer Query—1427

Writing & Preference Not comparable

Claude 2: —, Claude Haiku 4.5: 57.9 (#123)

Writing & Preference benchmarks
BenchmarkClaude 2Claude Haiku 4.5
LMArena Text—1396
LMArena Creative Writing—1372
WildBench—83.9%
EQ-Bench 4—1064
LMArena Multi-Turn—1409

Frequently asked questions

Is Claude 2 better than Claude Haiku 4.5?

Claude Haiku 4.5 is the stronger model overall, scoring 39.5 to 25.0 on the Noometry Index.

How many benchmarks do Claude 2 and Claude Haiku 4.5 share?

5 benchmarks have published results for both models. Claude 2 has 8 scored results on Noometry and Claude Haiku 4.5 has 53.

Related comparisons

Go deeper