Model comparison

Codellama 70b Instruct vs Llama 3.2 90B

Codellama 70b Instruct is the stronger model overall, scoring 33.7 to 27.5 on the Noometry Index.

Last verified . 0 shared benchmarks.

Codellama 70b Instruct Meta

33.7

Rank #237 Confirmed

Llama 3.2 90B Meta

27.5

Rank #331 Confirmed

Side by side

Codellama 70b Instruct and Llama 3.2 90B specifications
Codellama 70b InstructLlama 3.2 90B
ProviderMetaMeta
Noometry Index33.727.5
Released—2024-09-24
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked79

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Codellama 70b Instruct: 37.6 (#193), Llama 3.2 90B: —

Coding benchmarks
BenchmarkCodellama 70b InstructLlama 3.2 90B
BigCodeBench Instruct40.7%—
BigCodeBench Complete49.6%—
HumanEval+65.9%—

Agentic & Tool Use Not comparable

Codellama 70b Instruct: —, Llama 3.2 90B: 30.0 (#80)

Agentic & Tool Use benchmarks
BenchmarkCodellama 70b InstructLlama 3.2 90B
BALROG—27.3%

Reasoning Llama 3.2 90B leads

Codellama 70b Instruct: 20.1 (#242), Llama 3.2 90B: 21.7 (#217)

Reasoning benchmarks
BenchmarkCodellama 70b InstructLlama 3.2 90B
EnigmaEval—0.4%
LMArena Hard Prompts1052—
Epoch Capabilities Index—125.5

Math Not comparable

Codellama 70b Instruct: —, Llama 3.2 90B: 11.1 (#308)

Math benchmarks
BenchmarkCodellama 70b InstructLlama 3.2 90B
OTIS Mock AIME 2024-2025—2.6%
MATH Level 5—39.4%

Knowledge Not comparable

Codellama 70b Instruct: —, Llama 3.2 90B: 21.7 (#274)

Knowledge benchmarks
BenchmarkCodellama 70b InstructLlama 3.2 90B
GPQA Diamond—41%
MMLU—80.3%

Multimodal Not comparable

Codellama 70b Instruct: —, Llama 3.2 90B: 25.4 (#124)

Multimodal benchmarks
BenchmarkCodellama 70b InstructLlama 3.2 90B
LMArena Vision—1000
GeoBench—52%

Multilingual Not comparable

Codellama 70b Instruct: 24.8 (#288), Llama 3.2 90B: —

Multilingual benchmarks
BenchmarkCodellama 70b InstructLlama 3.2 90B
LMArena Non-English992—

Instruction Following Not comparable

Codellama 70b Instruct: 51.9 (#293), Llama 3.2 90B: —

Instruction Following benchmarks
BenchmarkCodellama 70b InstructLlama 3.2 90B
LMArena Instruction Following1024—

Writing & Preference Not comparable

Codellama 70b Instruct: 33.4 (#277), Llama 3.2 90B: —

Writing & Preference benchmarks
BenchmarkCodellama 70b InstructLlama 3.2 90B
LMArena Text1057—

Frequently asked questions

Is Codellama 70b Instruct better than Llama 3.2 90B?

Codellama 70b Instruct is the stronger model overall, scoring 33.7 to 27.5 on the Noometry Index.

How many benchmarks do Codellama 70b Instruct and Llama 3.2 90B share?

0 benchmarks have published results for both models. Codellama 70b Instruct has 7 scored results on Noometry and Llama 3.2 90B has 9.

Related comparisons

Go deeper