Model comparison

Codellama 70b Instruct vs Mistral Nemo

Codellama 70b Instruct is the stronger model overall, scoring 33.7 to 26.4 on the Noometry Index.

Last verified . 0 shared benchmarks.

Codellama 70b Instruct Meta

33.7

Rank #237 Confirmed

Mistral Nemo Mistral AI

26.4

Rank #337 Confirmed

Summary

  • The widest gap is in writing & preference, where Codellama 70b Instruct leads 33.4 to 28.5.

Side by side

Codellama 70b Instruct and Mistral Nemo specifications
Codellama 70b InstructMistral Nemo
ProviderMetaMistral AI
Noometry Index33.726.4
Released—2024-07-01
WeightsOpenOpen
Context window—128K
Max output—128K
Input $ / M tokens—$0.15
Output $ / M tokens—$0.15
Results tracked710

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Codellama 70b Instruct: 37.6 (#193), Mistral Nemo: —

Coding benchmarks
BenchmarkCodellama 70b InstructMistral Nemo
BigCodeBench Instruct40.7%—
BigCodeBench Complete49.6%—
HumanEval+65.9%—

Agentic & Tool Use Not comparable

Codellama 70b Instruct: —, Mistral Nemo: 23.5 (#125)

Agentic & Tool Use benchmarks
BenchmarkCodellama 70b InstructMistral Nemo
Berkeley Function Calling Leaderboard—27.6%
BALROG—17.6%

Reasoning Too close to call

Codellama 70b Instruct: 20.1 (#242), Mistral Nemo: 20.7 (#232)

Reasoning benchmarks
BenchmarkCodellama 70b InstructMistral Nemo
LMArena Hard Prompts1052—
DTBench—48.6%
Epoch Capabilities Index—118.68
PIQA—83.5%

Math Not comparable

Codellama 70b Instruct: —, Mistral Nemo: 25.5 (#268)

Math benchmarks
BenchmarkCodellama 70b InstructMistral Nemo
MATH Level 5—10.8%
GSM8K—84.2%

Knowledge Not comparable

Codellama 70b Instruct: —, Mistral Nemo: 12.3 (#298)

Knowledge benchmarks
BenchmarkCodellama 70b InstructMistral Nemo
GPQA Diamond—29.9%
BoolQ—82.5%

Multilingual Not comparable

Codellama 70b Instruct: 24.8 (#288), Mistral Nemo: —

Multilingual benchmarks
BenchmarkCodellama 70b InstructMistral Nemo
LMArena Non-English992—

Instruction Following Not comparable

Codellama 70b Instruct: 51.9 (#293), Mistral Nemo: —

Instruction Following benchmarks
BenchmarkCodellama 70b InstructMistral Nemo
LMArena Instruction Following1024—

Writing & Preference Codellama 70b Instruct leads

Codellama 70b Instruct: 33.4 (#277), Mistral Nemo: 28.5 (#296)

Writing & Preference benchmarks
BenchmarkCodellama 70b InstructMistral Nemo
LMArena Text1057—
EQ-Bench Creative Writing—881

Frequently asked questions

Is Codellama 70b Instruct better than Mistral Nemo?

Codellama 70b Instruct is the stronger model overall, scoring 33.7 to 26.4 on the Noometry Index.

How many benchmarks do Codellama 70b Instruct and Mistral Nemo share?

0 benchmarks have published results for both models. Codellama 70b Instruct has 7 scored results on Noometry and Mistral Nemo has 10.

Related comparisons

Go deeper