Model comparison

Codellama 34b Instruct vs Llama 3.2 90B

Codellama 34b Instruct is the stronger model overall, scoring 30.8 to 27.5 on the Noometry Index.

Last verified . 0 shared benchmarks.

Codellama 34b Instruct Meta

30.8

Rank #287 Confirmed

Llama 3.2 90B Meta

27.5

Rank #331 Confirmed

Summary

  • The widest gap is in math, where Codellama 34b Instruct leads 31.0 to 11.1.

Side by side

Codellama 34b Instruct and Llama 3.2 90B specifications
Codellama 34b InstructLlama 3.2 90B
ProviderMetaMeta
Noometry Index30.827.5
Released—2024-09-24
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked149

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Codellama 34b Instruct: 28.5 (#314), Llama 3.2 90B: —

Coding benchmarks
BenchmarkCodellama 34b InstructLlama 3.2 90B
BigCodeBench Instruct29%—
LMArena Coding1046—
BigCodeBench Complete37.1%—
HumanEval+43.9%—
MBPP+56.3%—

Agentic & Tool Use Not comparable

Codellama 34b Instruct: —, Llama 3.2 90B: 30.0 (#80)

Agentic & Tool Use benchmarks
BenchmarkCodellama 34b InstructLlama 3.2 90B
BALROG—27.3%

Reasoning Llama 3.2 90B leads

Codellama 34b Instruct: 19.6 (#255), Llama 3.2 90B: 21.7 (#217)

Reasoning benchmarks
BenchmarkCodellama 34b InstructLlama 3.2 90B
EnigmaEval—0.4%
LMArena Hard Prompts1032—
Epoch Capabilities Index—125.5

Math Codellama 34b Instruct leads

Codellama 34b Instruct: 31.0 (#230), Llama 3.2 90B: 11.1 (#308)

Math benchmarks
BenchmarkCodellama 34b InstructLlama 3.2 90B
OTIS Mock AIME 2024-2025—2.6%
LMArena Math1056—
MATH Level 5—39.4%

Knowledge Not comparable

Codellama 34b Instruct: —, Llama 3.2 90B: 21.7 (#274)

Knowledge benchmarks
BenchmarkCodellama 34b InstructLlama 3.2 90B
GPQA Diamond—41%
MMLU—80.3%

Multimodal Not comparable

Codellama 34b Instruct: —, Llama 3.2 90B: 25.4 (#124)

Multimodal benchmarks
BenchmarkCodellama 34b InstructLlama 3.2 90B
LMArena Vision—1000
GeoBench—52%

Multilingual Not comparable

Codellama 34b Instruct: 25.8 (#284), Llama 3.2 90B: —

Multilingual benchmarks
BenchmarkCodellama 34b InstructLlama 3.2 90B
LMArena Non-English1011—
LMArena Chinese976—

Instruction Following Not comparable

Codellama 34b Instruct: 52.2 (#291), Llama 3.2 90B: —

Instruction Following benchmarks
BenchmarkCodellama 34b InstructLlama 3.2 90B
LMArena Instruction Following1028—

Long Context Not comparable

Codellama 34b Instruct: 30.9 (#284), Llama 3.2 90B: —

Long Context benchmarks
BenchmarkCodellama 34b InstructLlama 3.2 90B
LMArena Longer Query1013—

Writing & Preference Not comparable

Codellama 34b Instruct: 28.2 (#297), Llama 3.2 90B: —

Writing & Preference benchmarks
BenchmarkCodellama 34b InstructLlama 3.2 90B
LMArena Text1066—
LMArena Creative Writing1032—
LMArena Multi-Turn1015—

Frequently asked questions

Is Codellama 34b Instruct better than Llama 3.2 90B?

Codellama 34b Instruct is the stronger model overall, scoring 30.8 to 27.5 on the Noometry Index.

How many benchmarks do Codellama 34b Instruct and Llama 3.2 90B share?

0 benchmarks have published results for both models. Codellama 34b Instruct has 14 scored results on Noometry and Llama 3.2 90B has 9.

Related comparisons

Go deeper