Model comparison

Claude 3.5 Haiku vs Codellama 34b Instruct

Codellama 34b Instruct is the stronger model overall, scoring 30.8 to 29.2 on the Noometry Index.

Last verified . 12 shared benchmarks.

Claude 3.5 Haiku Anthropic

29.2

Rank #315 Confirmed

Codellama 34b Instruct Meta

30.8

Rank #287 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Claude 3.5 Haiku scores higher in 5 categories and Codellama 34b Instruct in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Codellama 34b Instruct leads 31.0 to 14.7.
  • The biggest single-benchmark swing is BigCodeBench Complete: 59% for Claude 3.5 Haiku and 37.1% for Codellama 34b Instruct.
  • Codellama 34b Instruct has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Haiku and Codellama 34b Instruct specifications
Claude 3.5 HaikuCodellama 34b Instruct
ProviderAnthropicMeta
Noometry Index29.230.8
Released2024-10-22—
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4914

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.5 Haiku leads

Claude 3.5 Haiku: 32.9 (#265), Codellama 34b Instruct: 28.5 (#314)

Coding benchmarks
BenchmarkClaude 3.5 HaikuCodellama 34b Instruct
BigCodeBench Instruct46.1%29%
LMArena Coding12861046
BigCodeBench Complete59%37.1%
Aider Polyglot28%—
SciCode27.4%—
WeirdML30.7%—
LiveBench Coding51.4%—
CadEval32%—
HumanEval+—43.9%
MBPP+—56.3%

Agentic & Tool Use Not comparable

Claude 3.5 Haiku: 28.0 (#95), Codellama 34b Instruct: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 HaikuCodellama 34b Instruct
BALROG19.3%—

Reasoning Codellama 34b Instruct leads

Claude 3.5 Haiku: 17.7 (#290), Codellama 34b Instruct: 19.6 (#255)

Reasoning benchmarks
BenchmarkClaude 3.5 HaikuCodellama 34b Instruct
LMArena Hard Prompts12511032
CritPt0%—
LiveBench Reasoning28.1%—
DTBench56.7%—
LiveBench Data Analysis48.5%—
Epoch Capabilities Index127.15—
LiveBench43.5%—

Math Codellama 34b Instruct leads

Claude 3.5 Haiku: 14.7 (#300), Codellama 34b Instruct: 31.0 (#230)

Math benchmarks
BenchmarkClaude 3.5 HaikuCodellama 34b Instruct
LMArena Math12441056
OTIS Mock AIME 2024-20254.3%—
Omni-MATH22.4%—
LiveBench Math35.5%—
MATH Level 546.4%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Not comparable

Claude 3.5 Haiku: 18.7 (#281), Codellama 34b Instruct: —

Knowledge benchmarks
BenchmarkClaude 3.5 HaikuCodellama 34b Instruct
GPQA Diamond38.1%—
MMLU-Pro60.5%—
Confabulations36.7%—
GPQA (HELM)36.3%—
LMArena Expert1208—
MMLU74.3%—

Multimodal Not comparable

Claude 3.5 Haiku: 26.8 (#117), Codellama 34b Instruct: —

Multimodal benchmarks
BenchmarkClaude 3.5 HaikuCodellama 34b Instruct
LMArena Vision1092—
GeoBench34%—

Multilingual Claude 3.5 Haiku leads

Claude 3.5 Haiku: 40.0 (#218), Codellama 34b Instruct: 25.8 (#284)

Multilingual benchmarks
BenchmarkClaude 3.5 HaikuCodellama 34b Instruct
LMArena Non-English12381011
LMArena Chinese1229976
LMArena French1264—
LMArena German1237—
LMArena Japanese1175—
LMArena Korean1173—
LMArena Russian1253—
LMArena Spanish1261—

Instruction Following Claude 3.5 Haiku leads

Claude 3.5 Haiku: 62.9 (#234), Codellama 34b Instruct: 52.2 (#291)

Instruction Following benchmarks
BenchmarkClaude 3.5 HaikuCodellama 34b Instruct
LMArena Instruction Following12411028
LiveBench Instruction Following61.9%—
IFEval79.2%—

Long Context Claude 3.5 Haiku leads

Claude 3.5 Haiku: 38.3 (#200), Codellama 34b Instruct: 30.9 (#284)

Long Context benchmarks
BenchmarkClaude 3.5 HaikuCodellama 34b Instruct
LMArena Longer Query12611013

Writing & Preference Claude 3.5 Haiku leads

Claude 3.5 Haiku: 42.7 (#234), Codellama 34b Instruct: 28.2 (#297)

Writing & Preference benchmarks
BenchmarkClaude 3.5 HaikuCodellama 34b Instruct
LMArena Text12551066
LMArena Creative Writing12331032
LMArena Multi-Turn12651015
Short-Story Creative Writing73.5%—
EQ-Bench Creative Writing1146—
WildBench76%—
LiveBench Language35.4%—

Frequently asked questions

Is Claude 3.5 Haiku better than Codellama 34b Instruct?

Codellama 34b Instruct is the stronger model overall, scoring 30.8 to 29.2 on the Noometry Index.

Is Claude 3.5 Haiku or Codellama 34b Instruct better for coding?

Claude 3.5 Haiku scores higher on coding benchmarks: 32.9 versus 28.5 in the Noometry coding category.

How many benchmarks do Claude 3.5 Haiku and Codellama 34b Instruct share?

12 benchmarks have published results for both models. Claude 3.5 Haiku has 49 scored results on Noometry and Codellama 34b Instruct has 14.

Related comparisons

Go deeper