Model comparison

Claude 3.5 Haiku vs Claude 3 Opus

Claude 3.5 Haiku and Claude 3 Opus score almost the same on the Noometry Index (29.2 vs 29.5), so choose on price, context window or the category you care about most.

Last verified . 35 shared benchmarks.

Claude 3.5 Haiku Anthropic

29.2

Rank #315 Confirmed

Claude 3 Opus Anthropic

29.5

Rank #310 Confirmed

Summary

  • They share 35 benchmarks with published results for both. Claude 3.5 Haiku scores higher in 4 categories and Claude 3 Opus in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Claude 3 Opus leads 24.5 to 18.7.
  • The biggest single-benchmark swing is LiveBench Language: 35.4% for Claude 3.5 Haiku and 50.4% for Claude 3 Opus.

Side by side

Claude 3.5 Haiku and Claude 3 Opus specifications
Claude 3.5 HaikuClaude 3 Opus
ProviderAnthropicAnthropic
Noometry Index29.229.5
Released2024-10-222024-02-29
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4946

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Claude 3.5 Haiku: 32.9 (#265), Claude 3 Opus: 32.9 (#267)

Coding benchmarks
BenchmarkClaude 3.5 HaikuClaude 3 Opus
WeirdML30.7%19.2%
BigCodeBench Instruct46.1%45.5%
LiveBench Coding51.4%38.6%
LMArena Coding12861264
BigCodeBench Complete59%57.4%
Aider Polyglot28%—
SciCode27.4%—
CadEval32%—
HumanEval+—77.4%
MBPP+—73.3%

Agentic & Tool Use Claude 3.5 Haiku leads

Claude 3.5 Haiku: 28.0 (#95), Claude 3 Opus: 24.6 (#116)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 HaikuClaude 3 Opus
Cybench—10%
BALROG19.3%—
METR Time Horizons—29.5%

Reasoning Claude 3.5 Haiku leads

Claude 3.5 Haiku: 17.7 (#290), Claude 3 Opus: 14.6 (#324)

Reasoning benchmarks
BenchmarkClaude 3.5 HaikuClaude 3 Opus
LiveBench Reasoning28.1%40.6%
LMArena Hard Prompts12511245
DTBench56.7%61.6%
LiveBench Data Analysis48.5%57.9%
Epoch Capabilities Index127.15126.91
LiveBench43.5%49.2%
SimpleBench—23.5%
CritPt0%—
Chess Puzzles—5%
EnigmaEval—0.8%
LMCA—17%
ForecastBench—58.4
WinoGrande—88.5%

Math Too close to call

Claude 3.5 Haiku: 14.7 (#300), Claude 3 Opus: 14.8 (#299)

Math benchmarks
BenchmarkClaude 3.5 HaikuClaude 3 Opus
OTIS Mock AIME 2024-20254.3%4.7%
LiveBench Math35.5%43.6%
LMArena Math12441273
MATH Level 546.4%37.5%
Omni-MATH22.4%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Claude 3 Opus leads

Claude 3.5 Haiku: 18.7 (#281), Claude 3 Opus: 24.5 (#267)

Knowledge benchmarks
BenchmarkClaude 3.5 HaikuClaude 3 Opus
GPQA Diamond38.1%47.2%
Confabulations36.7%22.7%
LMArena Expert12081223
MMLU74.3%84.6%
SimpleQA Verified—12.6%
MMLU-Pro60.5%—
GPQA (HELM)36.3%—

Multimodal Too close to call

Claude 3.5 Haiku: 26.8 (#117), Claude 3 Opus: 27.1 (#116)

Multimodal benchmarks
BenchmarkClaude 3.5 HaikuClaude 3 Opus
LMArena Vision10921023
GeoBench34%—

Multilingual Claude 3 Opus leads

Claude 3.5 Haiku: 40.0 (#218), Claude 3 Opus: 41.4 (#207)

Multilingual benchmarks
BenchmarkClaude 3.5 HaikuClaude 3 Opus
LMArena Non-English12381258
LMArena Chinese12291248
LMArena French12641275
LMArena German12371258
LMArena Japanese11751204
LMArena Korean11731187
LMArena Russian12531280
LMArena Spanish12611246

Instruction Following Claude 3 Opus leads

Claude 3.5 Haiku: 62.9 (#234), Claude 3 Opus: 64.1 (#228)

Instruction Following benchmarks
BenchmarkClaude 3.5 HaikuClaude 3 Opus
LiveBench Instruction Following61.9%63.9%
LMArena Instruction Following12411248
IFEval79.2%—

Long Context Too close to call

Claude 3.5 Haiku: 38.3 (#200), Claude 3 Opus: 38.2 (#202)

Long Context benchmarks
BenchmarkClaude 3.5 HaikuClaude 3 Opus
LMArena Longer Query12611259

Writing & Preference Claude 3 Opus leads

Claude 3.5 Haiku: 42.7 (#234), Claude 3 Opus: 47.2 (#213)

Writing & Preference benchmarks
BenchmarkClaude 3.5 HaikuClaude 3 Opus
LMArena Text12551262
LMArena Creative Writing12331235
LMArena Multi-Turn12651275
LiveBench Language35.4%50.4%
Short-Story Creative Writing73.5%—
EQ-Bench Creative Writing1146—
WildBench76%—

Frequently asked questions

Is Claude 3.5 Haiku better than Claude 3 Opus?

Claude 3.5 Haiku and Claude 3 Opus score almost the same on the Noometry Index (29.2 vs 29.5), so choose on price, context window or the category you care about most.

Is Claude 3.5 Haiku or Claude 3 Opus better for coding?

They score almost the same on coding (32.9 vs 32.9); test both on your own repository before choosing.

How many benchmarks do Claude 3.5 Haiku and Claude 3 Opus share?

35 benchmarks have published results for both models. Claude 3.5 Haiku has 49 scored results on Noometry and Claude 3 Opus has 46.

Related comparisons

Go deeper