Model comparison

Claude 3 Opus vs Claude Haiku 4.5

Claude Haiku 4.5 is the stronger model overall, scoring 39.5 to 29.5 on the Noometry Index.

Last verified . 27 shared benchmarks.

Claude 3 Opus Anthropic

29.5

Rank #310 Confirmed

Claude Haiku 4.5 Anthropic

39.5

Rank #165 Confirmed

Summary

  • They share 27 benchmarks with published results for both. Claude 3 Opus scores higher in 1 category and Claude Haiku 4.5 in 9 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Haiku 4.5 leads 44.9 to 14.8.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 4.7% for Claude 3 Opus and 66.7% for Claude Haiku 4.5.

Side by side

Claude 3 Opus and Claude Haiku 4.5 specifications
Claude 3 OpusClaude Haiku 4.5
ProviderAnthropicAnthropic
Noometry Index29.539.5
Released2024-02-292025-10-15
WeightsProprietaryProprietary
Context window—200K
Max output—64K
Input $ / M tokens—$1
Output $ / M tokens—$5
Results tracked4653

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Haiku 4.5 leads

Claude 3 Opus: 32.9 (#267), Claude Haiku 4.5: 44.0 (#78)

Coding benchmarks
BenchmarkClaude 3 OpusClaude Haiku 4.5
WeirdML19.2%45.4%
LMArena Coding12641453
SWE-bench Verified (bash only)—66.6%
LMArena WebDev—1330
SWE-bench Multilingual—64.7%
SciCode—43.3%
BigCodeBench Instruct45.5%—
LiveBench Coding38.6%—
BigCodeBench Complete57.4%—
ALE-Bench—653.48
HumanEval+77.4%—
MBPP+73.3%—

Agentic & Tool Use Claude Haiku 4.5 leads

Claude 3 Opus: 24.6 (#116), Claude Haiku 4.5: 33.6 (#52)

Agentic & Tool Use benchmarks
BenchmarkClaude 3 OpusClaude Haiku 4.5
Terminal-Bench—35.5%
Berkeley Function Calling Leaderboard—68.7%
Cybench10%—
DeepResearch Bench—45.5%
BALROG—31.2%
ExploitBench—13.7%
METR Time Horizons29.5%—
Vending-Bench 2—458.89

Reasoning Too close to call

Claude 3 Opus: 14.6 (#324), Claude Haiku 4.5: 15.1 (#320)

Reasoning benchmarks
BenchmarkClaude 3 OpusClaude Haiku 4.5
Chess Puzzles5%8%
LMArena Hard Prompts12451420
DTBench61.6%73.6%
LMCA17%30.9%
Epoch Capabilities Index126.91142.41
ForecastBench58.461.4
ARC-AGI-2—4%
SimpleBench23.5%—
NYT Connections (extended)—14.3%
ARC-AGI-1—47.7%
CritPt—0%
EnigmaEval0.8%—
LiveBench Reasoning40.6%—
LiveBench Data Analysis57.9%—
LiveBench49.2%—
WinoGrande88.5%—

Math Claude Haiku 4.5 leads

Claude 3 Opus: 14.8 (#299), Claude Haiku 4.5: 44.9 (#78)

Math benchmarks
BenchmarkClaude 3 OpusClaude Haiku 4.5
OTIS Mock AIME 2024-20254.7%66.7%
LMArena Math12731396
MATH Level 537.5%96.4%
Omni-MATH—56.1%
LiveBench Math43.6%—
FrontierMath (Feb 2025 set)—5.9%
FrontierMath Tier 4 (v1)—2.1%

Knowledge Claude Haiku 4.5 leads

Claude 3 Opus: 24.5 (#267), Claude Haiku 4.5: 37.7 (#153)

Knowledge benchmarks
BenchmarkClaude 3 OpusClaude Haiku 4.5
GPQA Diamond47.2%71.2%
SimpleQA Verified12.6%13.2%
LMArena Expert12231442
MMLU-Pro—77.7%
Confabulations22.7%—
Vectara Hallucination Rate—9.8%
GPQA (HELM)—60.5%
MMLU84.6%—

Multimodal Too close to call

Claude 3 Opus: 27.1 (#116), Claude Haiku 4.5: 26.8 (#118)

Multimodal benchmarks
BenchmarkClaude 3 OpusClaude Haiku 4.5
LMArena Vision1023—
Blueprint-Bench 2—0%
LMArena Document—1420

Multilingual Claude Haiku 4.5 leads

Claude 3 Opus: 41.4 (#207), Claude Haiku 4.5: 49.9 (#129)

Multilingual benchmarks
BenchmarkClaude 3 OpusClaude Haiku 4.5
LMArena Non-English12581377
LMArena Chinese12481417
LMArena French12751408
LMArena German12581375
LMArena Japanese12041339
LMArena Korean11871347
LMArena Russian12801381
LMArena Spanish12461420

Instruction Following Claude Haiku 4.5 leads

Claude 3 Opus: 64.1 (#228), Claude Haiku 4.5: 71.4 (#149)

Instruction Following benchmarks
BenchmarkClaude 3 OpusClaude Haiku 4.5
LMArena Instruction Following12481414
LiveBench Instruction Following63.9%—
IFEval—80.1%

Long Context Claude Haiku 4.5 leads

Claude 3 Opus: 38.2 (#202), Claude Haiku 4.5: 43.6 (#92)

Long Context benchmarks
BenchmarkClaude 3 OpusClaude Haiku 4.5
LMArena Longer Query12591427

Writing & Preference Claude Haiku 4.5 leads

Claude 3 Opus: 47.2 (#213), Claude Haiku 4.5: 57.9 (#123)

Writing & Preference benchmarks
BenchmarkClaude 3 OpusClaude Haiku 4.5
LMArena Text12621396
LMArena Creative Writing12351372
LMArena Multi-Turn12751409
WildBench—83.9%
EQ-Bench 4—1064
LiveBench Language50.4%—

Frequently asked questions

Is Claude 3 Opus better than Claude Haiku 4.5?

Claude Haiku 4.5 is the stronger model overall, scoring 39.5 to 29.5 on the Noometry Index.

Is Claude 3 Opus or Claude Haiku 4.5 better for coding?

Claude Haiku 4.5 scores higher on coding benchmarks: 44.0 versus 32.9 in the Noometry coding category.

How many benchmarks do Claude 3 Opus and Claude Haiku 4.5 share?

27 benchmarks have published results for both models. Claude 3 Opus has 46 scored results on Noometry and Claude Haiku 4.5 has 53.

Related comparisons

Go deeper