Model comparison

Claude 3.5 Haiku vs Llama 3-70B

Claude 3.5 Haiku and Llama 3-70B score almost the same on the Noometry Index (29.2 vs 28.8), so choose on price, context window or the category you care about most.

Last verified . 25 shared benchmarks.

Claude 3.5 Haiku Anthropic

29.2

Rank #315 Confirmed

Llama 3-70B Meta

28.8

Rank #323 Confirmed

Summary

  • They share 25 benchmarks with published results for both. Claude 3.5 Haiku scores higher in 5 categories and Llama 3-70B in 4 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where Claude 3.5 Haiku leads 28.0 to 21.1.
  • The biggest single-benchmark swing is MATH Level 5: 46.4% for Claude 3.5 Haiku and 22.6% for Llama 3-70B.
  • Llama 3-70B has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Haiku and Llama 3-70B specifications
Claude 3.5 HaikuLlama 3-70B
ProviderAnthropicMeta
Noometry Index29.228.8
Released2024-10-222024-04-18
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4931

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3-70B leads

Claude 3.5 Haiku: 32.9 (#265), Llama 3-70B: 35.8 (#218)

Coding benchmarks
BenchmarkClaude 3.5 HaikuLlama 3-70B
BigCodeBench Instruct46.1%43.6%
LMArena Coding12861206
BigCodeBench Complete59%54.5%
Aider Polyglot28%—
SciCode27.4%—
WeirdML30.7%—
LiveBench Coding51.4%—
CadEval32%—
HumanEval+—72%
MBPP+—69%

Agentic & Tool Use Claude 3.5 Haiku leads

Claude 3.5 Haiku: 28.0 (#95), Llama 3-70B: 21.1 (#139)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 HaikuLlama 3-70B
Cybench—5%
BALROG19.3%—

Reasoning Too close to call

Claude 3.5 Haiku: 17.7 (#290), Llama 3-70B: 18.0 (#288)

Reasoning benchmarks
BenchmarkClaude 3.5 HaikuLlama 3-70B
LMArena Hard Prompts12511195
DTBench56.7%54.2%
Epoch Capabilities Index127.15122.93
Kagi LLM Benchmark—35.1%
CritPt0%—
LiveBench Reasoning28.1%—
LiveBench Data Analysis48.5%—
ForecastBench—57.1
LiveBench43.5%—
WinoGrande—83.5%

Math Claude 3.5 Haiku leads

Claude 3.5 Haiku: 14.7 (#300), Llama 3-70B: 12.8 (#305)

Math benchmarks
BenchmarkClaude 3.5 HaikuLlama 3-70B
OTIS Mock AIME 2024-20254.3%4.3%
LMArena Math12441218
MATH Level 546.4%22.6%
Omni-MATH22.4%—
LiveBench Math35.5%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Llama 3-70B leads

Claude 3.5 Haiku: 18.7 (#281), Llama 3-70B: 20.8 (#277)

Knowledge benchmarks
BenchmarkClaude 3.5 HaikuLlama 3-70B
GPQA Diamond38.1%40.6%
LMArena Expert12081149
MMLU74.3%79.3%
MMLU-Pro60.5%—
Confabulations36.7%—
GPQA (HELM)36.3%—

Multimodal Not comparable

Claude 3.5 Haiku: 26.8 (#117), Llama 3-70B: —

Multimodal benchmarks
BenchmarkClaude 3.5 HaikuLlama 3-70B
LMArena Vision1092—
GeoBench34%—

Multilingual Claude 3.5 Haiku leads

Claude 3.5 Haiku: 40.0 (#218), Llama 3-70B: 33.6 (#251)

Multilingual benchmarks
BenchmarkClaude 3.5 HaikuLlama 3-70B
LMArena Non-English12381142
LMArena Chinese12291114
LMArena French12641232
LMArena German12371169
LMArena Japanese11751017
LMArena Korean11731017
LMArena Russian12531159
LMArena Spanish12611241

Instruction Following Too close to call

Claude 3.5 Haiku: 62.9 (#234), Llama 3-70B: 62.5 (#238)

Instruction Following benchmarks
BenchmarkClaude 3.5 HaikuLlama 3-70B
LMArena Instruction Following12411194
LiveBench Instruction Following61.9%—
IFEval79.2%—

Long Context Claude 3.5 Haiku leads

Claude 3.5 Haiku: 38.3 (#200), Llama 3-70B: 35.6 (#240)

Long Context benchmarks
BenchmarkClaude 3.5 HaikuLlama 3-70B
LMArena Longer Query12611174

Writing & Preference Too close to call

Claude 3.5 Haiku: 42.7 (#234), Llama 3-70B: 42.8 (#231)

Writing & Preference benchmarks
BenchmarkClaude 3.5 HaikuLlama 3-70B
LMArena Text12551221
LMArena Creative Writing12331210
LMArena Multi-Turn12651223
Short-Story Creative Writing73.5%—
EQ-Bench Creative Writing1146—
WildBench76%—
LiveBench Language35.4%—

Frequently asked questions

Is Claude 3.5 Haiku better than Llama 3-70B?

Claude 3.5 Haiku and Llama 3-70B score almost the same on the Noometry Index (29.2 vs 28.8), so choose on price, context window or the category you care about most.

Is Claude 3.5 Haiku or Llama 3-70B better for coding?

Llama 3-70B scores higher on coding benchmarks: 35.8 versus 32.9 in the Noometry coding category.

How many benchmarks do Claude 3.5 Haiku and Llama 3-70B share?

25 benchmarks have published results for both models. Claude 3.5 Haiku has 49 scored results on Noometry and Llama 3-70B has 31.

Related comparisons

Go deeper