Model comparison

Devstral Small 2505 vs Mistral Large

Devstral Small 2505 is the stronger model overall, scoring 34.3 to 31.9 on the Noometry Index.

Last verified . 2 shared benchmarks.

Devstral Small 2505 Mistral AI

34.3

Rank #233 Reported

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Devstral Small 2505 scores higher in 2 categories and Mistral Large in 0 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Devstral Small 2505 leads 38.9 to 34.3.
  • The biggest single-benchmark swing is SciCode: 28.8% for Devstral Small 2505 and 36.2% for Mistral Large.
  • Devstral Small 2505 is cheaper at $0.10 / $0.30 per million input/output tokens, against $2 / $6 for Mistral Large.
  • Mistral Large accepts more context: 131K tokens versus 128K.

Side by side

Devstral Small 2505 and Mistral Large specifications
Devstral Small 2505Mistral Large
ProviderMistral AIMistral AI
Noometry Index34.331.9
Released2025-05-072024-02-26
WeightsOpenOpen
Context window128K131K
Max output128K16K
Input $ / M tokens$0.10$2
Output $ / M tokens$0.30$6
Results tracked451

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Devstral Small 2505 leads

Devstral Small 2505: 38.9 (#166), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkDevstral Small 2505Mistral Large
SciCode28.8%36.2%
SWE-bench Verified (bash only)56.4%—
BigCodeBench Instruct—30%
LiveBench Coding—47.1%
LMArena Coding—1277
BigCodeBench Complete—38.3%
ALE-Bench—264.7
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Not comparable

Devstral Small 2505: —, Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkDevstral Small 2505Mistral Large
Berkeley Function Calling Leaderboard—38.4%

Reasoning Devstral Small 2505 leads

Devstral Small 2505: 19.7 (#252), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkDevstral Small 2505Mistral Large
CritPt0%0%
SimpleBench—22.5%
Kagi LLM Benchmark37.7%—
LiveBench Reasoning—43.5%
LMArena Hard Prompts—1257
DTBench—65.1%
LiveBench Data Analysis—50.1%
LMCA—16.7%
Epoch Capabilities Index—128.52
ForecastBench—57.1
LiveBench—48.4%

Math Not comparable

Devstral Small 2505: —, Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkDevstral Small 2505Mistral Large
OTIS Mock AIME 2024-2025—8.5%
Omni-MATH—28.1%
LiveBench Math—42.5%
LMArena Math—1262
MATH Level 5—50.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Not comparable

Devstral Small 2505: —, Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkDevstral Small 2505Mistral Large
GPQA Diamond—51.3%
MMLU-Pro—59.9%
Confabulations—21.4%
Vectara Hallucination Rate—4.5%
GPQA (HELM)—43.5%
LMArena Expert—1232
MMLU—80%

Multilingual Not comparable

Devstral Small 2505: —, Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkDevstral Small 2505Mistral Large
LMArena Non-English—1237
LMArena Chinese—1240
LMArena French—1325
LMArena German—1254
LMArena Japanese—1188
LMArena Korean—1202
LMArena Russian—1257
LMArena Spanish—1268

Instruction Following Not comparable

Devstral Small 2505: —, Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkDevstral Small 2505Mistral Large
LiveBench Instruction Following—67.9%
IFEval—87.7%
LMArena Instruction Following—1249

Long Context Not comparable

Devstral Small 2505: —, Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkDevstral Small 2505Mistral Large
LMArena Longer Query—1261

Writing & Preference Not comparable

Devstral Small 2505: —, Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkDevstral Small 2505Mistral Large
LMArena Text—1266
LMArena Creative Writing—1243
Short-Story Creative Writing—69%
EQ-Bench Creative Writing—985
WildBench—80.1%
LMArena Multi-Turn—1260
LiveBench Language—39.4%

Frequently asked questions

Is Devstral Small 2505 better than Mistral Large?

Devstral Small 2505 is the stronger model overall, scoring 34.3 to 31.9 on the Noometry Index.

Which is cheaper, Devstral Small 2505 or Mistral Large?

Devstral Small 2505 is cheaper. It lists at $0.10 per million input tokens and $0.30 per million output tokens; Mistral Large lists at $2 and $6.

Is Devstral Small 2505 or Mistral Large better for coding?

Devstral Small 2505 scores higher on coding benchmarks: 38.9 versus 34.3 in the Noometry coding category.

Which has the bigger context window?

Mistral Large does, with 131K tokens against 128K.

How many benchmarks do Devstral Small 2505 and Mistral Large share?

2 benchmarks have published results for both models. Devstral Small 2505 has 4 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper