Model comparison

Codellama 34b Instruct vs Llama 13b

Codellama 34b Instruct is the stronger model overall, scoring 30.8 to 24.4 on the Noometry Index.

Last verified . 8 shared benchmarks.

Codellama 34b Instruct Meta

30.8

Rank #287 Confirmed

Llama 13b Meta

24.4

Rank #348 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Codellama 34b Instruct scores higher in 6 categories and Llama 13b in 0 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Codellama 34b Instruct leads 52.2 to 36.7.

Side by side

Codellama 34b Instruct and Llama 13b specifications
Codellama 34b InstructLlama 13b
ProviderMetaMeta
Noometry Index30.824.4
Released—2023-02-24
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1421

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Codellama 34b Instruct leads

Codellama 34b Instruct: 28.5 (#314), Llama 13b: 21.4 (#337)

Coding benchmarks
BenchmarkCodellama 34b InstructLlama 13b
LMArena Coding1046683
BigCodeBench Instruct29%—
BigCodeBench Complete37.1%—
HumanEval+43.9%—
MBPP+56.3%—

Reasoning Codellama 34b Instruct leads

Codellama 34b Instruct: 19.6 (#255), Llama 13b: 14.0 (#329)

Reasoning benchmarks
BenchmarkCodellama 34b InstructLlama 13b
LMArena Hard Prompts1032728
BIG-Bench Hard—37.9%
Epoch Capabilities Index—100.58
HellaSwag—79.2%
LAMBADA—75.2%
PIQA—80.1%
WinoGrande—73%

Math Codellama 34b Instruct leads

Codellama 34b Instruct: 31.0 (#230), Llama 13b: 26.7 (#256)

Math benchmarks
BenchmarkCodellama 34b InstructLlama 13b
LMArena Math1056838
GSM8K—20.6%

Knowledge Not comparable

Codellama 34b Instruct: —, Llama 13b: —

Knowledge benchmarks
BenchmarkCodellama 34b InstructLlama 13b
ARC (AI2) Challenge—52.7%
BoolQ—78.7%
MMLU—47.7%
OpenBookQA—56.4%
TriviaQA—77.9%

Multimodal Not comparable

Codellama 34b Instruct: —, Llama 13b: —

Multimodal benchmarks
BenchmarkCodellama 34b InstructLlama 13b
ScienceQA—43.3%

Multilingual Codellama 34b Instruct leads

Codellama 34b Instruct: 25.8 (#284), Llama 13b: 16.6 (#297)

Multilingual benchmarks
BenchmarkCodellama 34b InstructLlama 13b
LMArena Non-English1011819
LMArena Chinese976—

Instruction Following Codellama 34b Instruct leads

Codellama 34b Instruct: 52.2 (#291), Llama 13b: 36.7 (#305)

Instruction Following benchmarks
BenchmarkCodellama 34b InstructLlama 13b
LMArena Instruction Following1028781

Long Context Not comparable

Codellama 34b Instruct: 30.9 (#284), Llama 13b: —

Long Context benchmarks
BenchmarkCodellama 34b InstructLlama 13b
LMArena Longer Query1013—

Writing & Preference Codellama 34b Instruct leads

Codellama 34b Instruct: 28.2 (#297), Llama 13b: 13.8 (#312)

Writing & Preference benchmarks
BenchmarkCodellama 34b InstructLlama 13b
LMArena Text1066834
LMArena Creative Writing1032794
LMArena Multi-Turn1015753

Frequently asked questions

Is Codellama 34b Instruct better than Llama 13b?

Codellama 34b Instruct is the stronger model overall, scoring 30.8 to 24.4 on the Noometry Index.

Is Codellama 34b Instruct or Llama 13b better for coding?

Codellama 34b Instruct scores higher on coding benchmarks: 28.5 versus 21.4 in the Noometry coding category.

How many benchmarks do Codellama 34b Instruct and Llama 13b share?

8 benchmarks have published results for both models. Codellama 34b Instruct has 14 scored results on Noometry and Llama 13b has 21.

Related comparisons

Go deeper