Model comparison

Codellama 34b Instruct vs Mistral Large

Mistral Large is the stronger model overall, scoring 31.9 to 30.8 on the Noometry Index.

Last verified . 14 shared benchmarks.

Codellama 34b Instruct Meta

30.8

Rank #287 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Codellama 34b Instruct scores higher in 2 categories and Mistral Large in 5 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Mistral Large leads 67.9 to 52.2.

Side by side

Codellama 34b Instruct and Mistral Large specifications
Codellama 34b InstructMistral Large
ProviderMetaMistral AI
Noometry Index30.831.9
Released—2024-02-26
WeightsOpenOpen
Context window—131K
Max output—16K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked1451

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Large leads

Codellama 34b Instruct: 28.5 (#314), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkCodellama 34b InstructMistral Large
BigCodeBench Instruct29%30%
LMArena Coding10461277
BigCodeBench Complete37.1%38.3%
HumanEval+43.9%62.2%
MBPP+56.3%59.5%
SciCode—36.2%
LiveBench Coding—47.1%
ALE-Bench—264.7

Agentic & Tool Use Not comparable

Codellama 34b Instruct: —, Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkCodellama 34b InstructMistral Large
Berkeley Function Calling Leaderboard—38.4%

Reasoning Codellama 34b Instruct leads

Codellama 34b Instruct: 19.6 (#255), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkCodellama 34b InstructMistral Large
LMArena Hard Prompts10321257
SimpleBench—22.5%
CritPt—0%
LiveBench Reasoning—43.5%
DTBench—65.1%
LiveBench Data Analysis—50.1%
LMCA—16.7%
Epoch Capabilities Index—128.52
ForecastBench—57.1
LiveBench—48.4%

Math Codellama 34b Instruct leads

Codellama 34b Instruct: 31.0 (#230), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkCodellama 34b InstructMistral Large
LMArena Math10561262
OTIS Mock AIME 2024-2025—8.5%
Omni-MATH—28.1%
LiveBench Math—42.5%
MATH Level 5—50.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Not comparable

Codellama 34b Instruct: —, Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkCodellama 34b InstructMistral Large
GPQA Diamond—51.3%
MMLU-Pro—59.9%
Confabulations—21.4%
Vectara Hallucination Rate—4.5%
GPQA (HELM)—43.5%
LMArena Expert—1232
MMLU—80%

Multilingual Mistral Large leads

Codellama 34b Instruct: 25.8 (#284), Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkCodellama 34b InstructMistral Large
LMArena Non-English10111237
LMArena Chinese9761240
LMArena French—1325
LMArena German—1254
LMArena Japanese—1188
LMArena Korean—1202
LMArena Russian—1257
LMArena Spanish—1268

Instruction Following Mistral Large leads

Codellama 34b Instruct: 52.2 (#291), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkCodellama 34b InstructMistral Large
LMArena Instruction Following10281249
LiveBench Instruction Following—67.9%
IFEval—87.7%

Long Context Mistral Large leads

Codellama 34b Instruct: 30.9 (#284), Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkCodellama 34b InstructMistral Large
LMArena Longer Query10131261

Writing & Preference Mistral Large leads

Codellama 34b Instruct: 28.2 (#297), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkCodellama 34b InstructMistral Large
LMArena Text10661266
LMArena Creative Writing10321243
LMArena Multi-Turn10151260
Short-Story Creative Writing—69%
EQ-Bench Creative Writing—985
WildBench—80.1%
LiveBench Language—39.4%

Frequently asked questions

Is Codellama 34b Instruct better than Mistral Large?

Mistral Large is the stronger model overall, scoring 31.9 to 30.8 on the Noometry Index.

Is Codellama 34b Instruct or Mistral Large better for coding?

Mistral Large scores higher on coding benchmarks: 34.3 versus 28.5 in the Noometry coding category.

How many benchmarks do Codellama 34b Instruct and Mistral Large share?

14 benchmarks have published results for both models. Codellama 34b Instruct has 14 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper