Model comparison

Codellama 34b Instruct vs Devstral Small 2505

Devstral Small 2505 is the stronger model overall, scoring 34.3 to 30.8 on the Noometry Index.

Last verified . 0 shared benchmarks.

Codellama 34b Instruct Meta

30.8

Rank #287 Confirmed

Devstral Small 2505 Mistral AI

34.3

Rank #233 Reported

Summary

  • The widest gap is in coding, where Devstral Small 2505 leads 38.9 to 28.5.

Side by side

Codellama 34b Instruct and Devstral Small 2505 specifications
Codellama 34b InstructDevstral Small 2505
ProviderMetaMistral AI
Noometry Index30.834.3
Released—2025-05-07
WeightsOpenOpen
Context window—128K
Max output—128K
Input $ / M tokens—$0.10
Output $ / M tokens—$0.30
Results tracked144

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Devstral Small 2505 leads

Codellama 34b Instruct: 28.5 (#314), Devstral Small 2505: 38.9 (#166)

Coding benchmarks
BenchmarkCodellama 34b InstructDevstral Small 2505
SWE-bench Verified (bash only)—56.4%
SciCode—28.8%
BigCodeBench Instruct29%—
LMArena Coding1046—
BigCodeBench Complete37.1%—
HumanEval+43.9%—
MBPP+56.3%—

Reasoning Too close to call

Codellama 34b Instruct: 19.6 (#255), Devstral Small 2505: 19.7 (#252)

Reasoning benchmarks
BenchmarkCodellama 34b InstructDevstral Small 2505
Kagi LLM Benchmark—37.7%
CritPt—0%
LMArena Hard Prompts1032—

Math Not comparable

Codellama 34b Instruct: 31.0 (#230), Devstral Small 2505: —

Math benchmarks
BenchmarkCodellama 34b InstructDevstral Small 2505
LMArena Math1056—

Multilingual Not comparable

Codellama 34b Instruct: 25.8 (#284), Devstral Small 2505: —

Multilingual benchmarks
BenchmarkCodellama 34b InstructDevstral Small 2505
LMArena Non-English1011—
LMArena Chinese976—

Instruction Following Not comparable

Codellama 34b Instruct: 52.2 (#291), Devstral Small 2505: —

Instruction Following benchmarks
BenchmarkCodellama 34b InstructDevstral Small 2505
LMArena Instruction Following1028—

Long Context Not comparable

Codellama 34b Instruct: 30.9 (#284), Devstral Small 2505: —

Long Context benchmarks
BenchmarkCodellama 34b InstructDevstral Small 2505
LMArena Longer Query1013—

Writing & Preference Not comparable

Codellama 34b Instruct: 28.2 (#297), Devstral Small 2505: —

Writing & Preference benchmarks
BenchmarkCodellama 34b InstructDevstral Small 2505
LMArena Text1066—
LMArena Creative Writing1032—
LMArena Multi-Turn1015—

Frequently asked questions

Is Codellama 34b Instruct better than Devstral Small 2505?

Devstral Small 2505 is the stronger model overall, scoring 34.3 to 30.8 on the Noometry Index.

Is Codellama 34b Instruct or Devstral Small 2505 better for coding?

Devstral Small 2505 scores higher on coding benchmarks: 38.9 versus 28.5 in the Noometry coding category.

How many benchmarks do Codellama 34b Instruct and Devstral Small 2505 share?

0 benchmarks have published results for both models. Codellama 34b Instruct has 14 scored results on Noometry and Devstral Small 2505 has 4.

Related comparisons

Go deeper