Model comparison

Codellama 34b Instruct vs Llama 3.1 Tulu 3 8b

Llama 3.1 Tulu 3 8b is the stronger model overall, scoring 35.7 to 30.8 on the Noometry Index.

Last verified . 10 shared benchmarks.

Codellama 34b Instruct Meta

30.8

Rank #287 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Codellama 34b Instruct scores higher in 0 categories and Llama 3.1 Tulu 3 8b in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Llama 3.1 Tulu 3 8b leads 39.7 to 28.2.

Side by side

Codellama 34b Instruct and Llama 3.1 Tulu 3 8b specifications
Codellama 34b InstructLlama 3.1 Tulu 3 8b
ProviderMetaAllen Institute for AI (Ai2)
Noometry Index30.835.7
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1411

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.1 Tulu 3 8b leads

Codellama 34b Instruct: 28.5 (#314), Llama 3.1 Tulu 3 8b: 34.4 (#235)

Coding benchmarks
BenchmarkCodellama 34b InstructLlama 3.1 Tulu 3 8b
LMArena Coding10461183
BigCodeBench Instruct29%—
BigCodeBench Complete37.1%—
HumanEval+43.9%—
MBPP+56.3%—

Reasoning Llama 3.1 Tulu 3 8b leads

Codellama 34b Instruct: 19.6 (#255), Llama 3.1 Tulu 3 8b: 22.8 (#188)

Reasoning benchmarks
BenchmarkCodellama 34b InstructLlama 3.1 Tulu 3 8b
LMArena Hard Prompts10321174

Math Llama 3.1 Tulu 3 8b leads

Codellama 34b Instruct: 31.0 (#230), Llama 3.1 Tulu 3 8b: 33.9 (#198)

Math benchmarks
BenchmarkCodellama 34b InstructLlama 3.1 Tulu 3 8b
LMArena Math10561195

Multilingual Llama 3.1 Tulu 3 8b leads

Codellama 34b Instruct: 25.8 (#284), Llama 3.1 Tulu 3 8b: 35.4 (#246)

Multilingual benchmarks
BenchmarkCodellama 34b InstructLlama 3.1 Tulu 3 8b
LMArena Non-English10111169
LMArena Chinese9761176
LMArena Russian—1193

Instruction Following Llama 3.1 Tulu 3 8b leads

Codellama 34b Instruct: 52.2 (#291), Llama 3.1 Tulu 3 8b: 61.3 (#246)

Instruction Following benchmarks
BenchmarkCodellama 34b InstructLlama 3.1 Tulu 3 8b
LMArena Instruction Following10281174

Long Context Llama 3.1 Tulu 3 8b leads

Codellama 34b Instruct: 30.9 (#284), Llama 3.1 Tulu 3 8b: 35.8 (#239)

Long Context benchmarks
BenchmarkCodellama 34b InstructLlama 3.1 Tulu 3 8b
LMArena Longer Query10131181

Writing & Preference Llama 3.1 Tulu 3 8b leads

Codellama 34b Instruct: 28.2 (#297), Llama 3.1 Tulu 3 8b: 39.7 (#245)

Writing & Preference benchmarks
BenchmarkCodellama 34b InstructLlama 3.1 Tulu 3 8b
LMArena Text10661193
LMArena Creative Writing10321182
LMArena Multi-Turn10151154

Frequently asked questions

Is Codellama 34b Instruct better than Llama 3.1 Tulu 3 8b?

Llama 3.1 Tulu 3 8b is the stronger model overall, scoring 35.7 to 30.8 on the Noometry Index.

Is Codellama 34b Instruct or Llama 3.1 Tulu 3 8b better for coding?

Llama 3.1 Tulu 3 8b scores higher on coding benchmarks: 34.4 versus 28.5 in the Noometry coding category.

How many benchmarks do Codellama 34b Instruct and Llama 3.1 Tulu 3 8b share?

10 benchmarks have published results for both models. Codellama 34b Instruct has 14 scored results on Noometry and Llama 3.1 Tulu 3 8b has 11.

Related comparisons

Go deeper