Model comparison

Codellama 70b Instruct vs Mistral

Codellama 70b Instruct is the stronger model overall, scoring 33.7 to 29.9 on the Noometry Index.

Last verified . 4 shared benchmarks.

Codellama 70b Instruct Meta

33.7

Rank #237 Confirmed

Mistral Mistral AI

29.9

Rank #303 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Codellama 70b Instruct scores higher in 1 category and Mistral in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Mistral leads 32.8 to 24.8.
  • Codellama 70b Instruct has downloadable open weights; the other is API-only.

Side by side

Codellama 70b Instruct and Mistral specifications
Codellama 70b InstructMistral
ProviderMetaMistral AI
Noometry Index33.729.9
Released——
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked722

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Codellama 70b Instruct leads

Codellama 70b Instruct: 37.6 (#193), Mistral: 33.8 (#250)

Coding benchmarks
BenchmarkCodellama 70b InstructMistral
BigCodeBench Instruct40.7%—
LMArena Coding—1162
BigCodeBench Complete49.6%—
HumanEval+65.9%—

Reasoning Mistral leads

Codellama 70b Instruct: 20.1 (#242), Mistral: 22.2 (#200)

Reasoning benchmarks
BenchmarkCodellama 70b InstructMistral
LMArena Hard Prompts10521149

Math Not comparable

Codellama 70b Instruct: —, Mistral: 22.3 (#278)

Math benchmarks
BenchmarkCodellama 70b InstructMistral
Omni-MATH—7.2%
LMArena Math—1180

Knowledge Not comparable

Codellama 70b Instruct: —, Mistral: 16.6 (#288)

Knowledge benchmarks
BenchmarkCodellama 70b InstructMistral
MMLU-Pro—27.7%
GPQA (HELM)—30.3%
LMArena Expert—1125

Multilingual Mistral leads

Codellama 70b Instruct: 24.8 (#288), Mistral: 32.8 (#254)

Multilingual benchmarks
BenchmarkCodellama 70b InstructMistral
LMArena Non-English9921129
LMArena Chinese—1109
LMArena French—1180
LMArena German—1155
LMArena Japanese—1013
LMArena Korean—1032
LMArena Russian—1168
LMArena Spanish—1143

Instruction Following Too close to call

Codellama 70b Instruct: 51.9 (#293), Mistral: 52.6 (#288)

Instruction Following benchmarks
BenchmarkCodellama 70b InstructMistral
LMArena Instruction Following10241152
IFEval—56.8%

Long Context Not comparable

Codellama 70b Instruct: —, Mistral: 35.0 (#245)

Long Context benchmarks
BenchmarkCodellama 70b InstructMistral
LMArena Longer Query—1153

Writing & Preference Mistral leads

Codellama 70b Instruct: 33.4 (#277), Mistral: 37.0 (#260)

Writing & Preference benchmarks
BenchmarkCodellama 70b InstructMistral
LMArena Text10571165
LMArena Creative Writing—1158
WildBench—66%
LMArena Multi-Turn—1147

Frequently asked questions

Is Codellama 70b Instruct better than Mistral?

Codellama 70b Instruct is the stronger model overall, scoring 33.7 to 29.9 on the Noometry Index.

Is Codellama 70b Instruct or Mistral better for coding?

Codellama 70b Instruct scores higher on coding benchmarks: 37.6 versus 33.8 in the Noometry coding category.

How many benchmarks do Codellama 70b Instruct and Mistral share?

4 benchmarks have published results for both models. Codellama 70b Instruct has 7 scored results on Noometry and Mistral has 22.

Related comparisons

Go deeper