Model comparison

Claude 3.5 Haiku vs Codellama 70b Instruct

Codellama 70b Instruct is the stronger model overall, scoring 33.7 to 29.2 on the Noometry Index.

Last verified . 6 shared benchmarks.

Claude 3.5 Haiku Anthropic

29.2

Rank #315 Confirmed

Codellama 70b Instruct Meta

33.7

Rank #237 Confirmed

Summary

  • They share 6 benchmarks with published results for both. Claude 3.5 Haiku scores higher in 3 categories and Codellama 70b Instruct in 2 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Claude 3.5 Haiku leads 40.0 to 24.8.
  • The biggest single-benchmark swing is BigCodeBench Complete: 59% for Claude 3.5 Haiku and 49.6% for Codellama 70b Instruct.
  • Codellama 70b Instruct has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Haiku and Codellama 70b Instruct specifications
Claude 3.5 HaikuCodellama 70b Instruct
ProviderAnthropicMeta
Noometry Index29.233.7
Released2024-10-22—
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked497

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Codellama 70b Instruct leads

Claude 3.5 Haiku: 32.9 (#265), Codellama 70b Instruct: 37.6 (#193)

Coding benchmarks
BenchmarkClaude 3.5 HaikuCodellama 70b Instruct
BigCodeBench Instruct46.1%40.7%
BigCodeBench Complete59%49.6%
Aider Polyglot28%—
SciCode27.4%—
WeirdML30.7%—
LiveBench Coding51.4%—
LMArena Coding1286—
CadEval32%—
HumanEval+—65.9%

Agentic & Tool Use Not comparable

Claude 3.5 Haiku: 28.0 (#95), Codellama 70b Instruct: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 HaikuCodellama 70b Instruct
BALROG19.3%—

Reasoning Codellama 70b Instruct leads

Claude 3.5 Haiku: 17.7 (#290), Codellama 70b Instruct: 20.1 (#242)

Reasoning benchmarks
BenchmarkClaude 3.5 HaikuCodellama 70b Instruct
LMArena Hard Prompts12511052
CritPt0%—
LiveBench Reasoning28.1%—
DTBench56.7%—
LiveBench Data Analysis48.5%—
Epoch Capabilities Index127.15—
LiveBench43.5%—

Math Not comparable

Claude 3.5 Haiku: 14.7 (#300), Codellama 70b Instruct: —

Math benchmarks
BenchmarkClaude 3.5 HaikuCodellama 70b Instruct
OTIS Mock AIME 2024-20254.3%—
Omni-MATH22.4%—
LiveBench Math35.5%—
LMArena Math1244—
MATH Level 546.4%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Not comparable

Claude 3.5 Haiku: 18.7 (#281), Codellama 70b Instruct: —

Knowledge benchmarks
BenchmarkClaude 3.5 HaikuCodellama 70b Instruct
GPQA Diamond38.1%—
MMLU-Pro60.5%—
Confabulations36.7%—
GPQA (HELM)36.3%—
LMArena Expert1208—
MMLU74.3%—

Multimodal Not comparable

Claude 3.5 Haiku: 26.8 (#117), Codellama 70b Instruct: —

Multimodal benchmarks
BenchmarkClaude 3.5 HaikuCodellama 70b Instruct
LMArena Vision1092—
GeoBench34%—

Multilingual Claude 3.5 Haiku leads

Claude 3.5 Haiku: 40.0 (#218), Codellama 70b Instruct: 24.8 (#288)

Multilingual benchmarks
BenchmarkClaude 3.5 HaikuCodellama 70b Instruct
LMArena Non-English1238992
LMArena Chinese1229—
LMArena French1264—
LMArena German1237—
LMArena Japanese1175—
LMArena Korean1173—
LMArena Russian1253—
LMArena Spanish1261—

Instruction Following Claude 3.5 Haiku leads

Claude 3.5 Haiku: 62.9 (#234), Codellama 70b Instruct: 51.9 (#293)

Instruction Following benchmarks
BenchmarkClaude 3.5 HaikuCodellama 70b Instruct
LMArena Instruction Following12411024
LiveBench Instruction Following61.9%—
IFEval79.2%—

Long Context Not comparable

Claude 3.5 Haiku: 38.3 (#200), Codellama 70b Instruct: —

Long Context benchmarks
BenchmarkClaude 3.5 HaikuCodellama 70b Instruct
LMArena Longer Query1261—

Writing & Preference Claude 3.5 Haiku leads

Claude 3.5 Haiku: 42.7 (#234), Codellama 70b Instruct: 33.4 (#277)

Writing & Preference benchmarks
BenchmarkClaude 3.5 HaikuCodellama 70b Instruct
LMArena Text12551057
LMArena Creative Writing1233—
Short-Story Creative Writing73.5%—
EQ-Bench Creative Writing1146—
WildBench76%—
LMArena Multi-Turn1265—
LiveBench Language35.4%—

Frequently asked questions

Is Claude 3.5 Haiku better than Codellama 70b Instruct?

Codellama 70b Instruct is the stronger model overall, scoring 33.7 to 29.2 on the Noometry Index.

Is Claude 3.5 Haiku or Codellama 70b Instruct better for coding?

Codellama 70b Instruct scores higher on coding benchmarks: 37.6 versus 32.9 in the Noometry coding category.

How many benchmarks do Claude 3.5 Haiku and Codellama 70b Instruct share?

6 benchmarks have published results for both models. Claude 3.5 Haiku has 49 scored results on Noometry and Codellama 70b Instruct has 7.

Related comparisons

Go deeper