Model comparison

Codellama 70b Instruct vs Mistral Large

Codellama 70b Instruct is the stronger model overall, scoring 33.7 to 31.9 on the Noometry Index.

Last verified . 7 shared benchmarks.

Codellama 70b Instruct Meta

33.7

Rank #237 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 7 benchmarks with published results for both. Codellama 70b Instruct scores higher in 2 categories and Mistral Large in 3 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Mistral Large leads 67.9 to 51.9.
  • The biggest single-benchmark swing is BigCodeBench Complete: 49.6% for Codellama 70b Instruct and 38.3% for Mistral Large.

Side by side

Codellama 70b Instruct and Mistral Large specifications
Codellama 70b InstructMistral Large
ProviderMetaMistral AI
Noometry Index33.731.9
Released—2024-02-26
WeightsOpenOpen
Context window—131K
Max output—16K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked751

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Codellama 70b Instruct leads

Codellama 70b Instruct: 37.6 (#193), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkCodellama 70b InstructMistral Large
BigCodeBench Instruct40.7%30%
BigCodeBench Complete49.6%38.3%
HumanEval+65.9%62.2%
SciCode—36.2%
LiveBench Coding—47.1%
LMArena Coding—1277
ALE-Bench—264.7
MBPP+—59.5%

Agentic & Tool Use Not comparable

Codellama 70b Instruct: —, Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkCodellama 70b InstructMistral Large
Berkeley Function Calling Leaderboard—38.4%

Reasoning Codellama 70b Instruct leads

Codellama 70b Instruct: 20.1 (#242), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkCodellama 70b InstructMistral Large
LMArena Hard Prompts10521257
SimpleBench—22.5%
CritPt—0%
LiveBench Reasoning—43.5%
DTBench—65.1%
LiveBench Data Analysis—50.1%
LMCA—16.7%
Epoch Capabilities Index—128.52
ForecastBench—57.1
LiveBench—48.4%

Math Not comparable

Codellama 70b Instruct: —, Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkCodellama 70b InstructMistral Large
OTIS Mock AIME 2024-2025—8.5%
Omni-MATH—28.1%
LiveBench Math—42.5%
LMArena Math—1262
MATH Level 5—50.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Not comparable

Codellama 70b Instruct: —, Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkCodellama 70b InstructMistral Large
GPQA Diamond—51.3%
MMLU-Pro—59.9%
Confabulations—21.4%
Vectara Hallucination Rate—4.5%
GPQA (HELM)—43.5%
LMArena Expert—1232
MMLU—80%

Multilingual Mistral Large leads

Codellama 70b Instruct: 24.8 (#288), Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkCodellama 70b InstructMistral Large
LMArena Non-English9921237
LMArena Chinese—1240
LMArena French—1325
LMArena German—1254
LMArena Japanese—1188
LMArena Korean—1202
LMArena Russian—1257
LMArena Spanish—1268

Instruction Following Mistral Large leads

Codellama 70b Instruct: 51.9 (#293), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkCodellama 70b InstructMistral Large
LMArena Instruction Following10241249
LiveBench Instruction Following—67.9%
IFEval—87.7%

Long Context Not comparable

Codellama 70b Instruct: —, Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkCodellama 70b InstructMistral Large
LMArena Longer Query—1261

Writing & Preference Mistral Large leads

Codellama 70b Instruct: 33.4 (#277), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkCodellama 70b InstructMistral Large
LMArena Text10571266
LMArena Creative Writing—1243
Short-Story Creative Writing—69%
EQ-Bench Creative Writing—985
WildBench—80.1%
LMArena Multi-Turn—1260
LiveBench Language—39.4%

Frequently asked questions

Is Codellama 70b Instruct better than Mistral Large?

Codellama 70b Instruct is the stronger model overall, scoring 33.7 to 31.9 on the Noometry Index.

Is Codellama 70b Instruct or Mistral Large better for coding?

Codellama 70b Instruct scores higher on coding benchmarks: 37.6 versus 34.3 in the Noometry coding category.

How many benchmarks do Codellama 70b Instruct and Mistral Large share?

7 benchmarks have published results for both models. Codellama 70b Instruct has 7 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper