Model comparison

Devstral Small 2505 vs Mercury

Mercury is the stronger model overall, scoring 37.6 to 34.3 on the Noometry Index.

Last verified . 1 shared benchmarks.

Devstral Small 2505 Mistral AI

34.3

Rank #233 Reported

Mercury Inception

37.6

Rank #199 Confirmed

Summary

  • They share 1 benchmark with published results for both. Devstral Small 2505 scores higher in 2 categories and Mercury in 0 categories; one gap is clear of the uncertainty.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 37.7% for Devstral Small 2505 and 21.6% for Mercury.
  • Devstral Small 2505 has downloadable open weights; the other is API-only.

Side by side

Devstral Small 2505 and Mercury specifications
Devstral Small 2505Mercury
ProviderMistral AIInception
Noometry Index34.337.6
Released2025-05-07—
WeightsOpenProprietary
Context window128K—
Max output128K—
Input $ / M tokens$0.10—
Output $ / M tokens$0.30—
Results tracked49

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Devstral Small 2505: 38.9 (#166), Mercury: 38.7 (#170)

Coding benchmarks
BenchmarkDevstral Small 2505Mercury
SWE-bench Verified (bash only)56.4%—
SciCode28.8%—
LMArena Coding—1322

Reasoning Devstral Small 2505 leads

Devstral Small 2505: 19.7 (#252), Mercury: 17.5 (#293)

Reasoning benchmarks
BenchmarkDevstral Small 2505Mercury
Kagi LLM Benchmark37.7%21.6%
CritPt0%—
LMArena Hard Prompts—1285

Multilingual Not comparable

Devstral Small 2505: —, Mercury: 41.6 (#206)

Multilingual benchmarks
BenchmarkDevstral Small 2505Mercury
LMArena Non-English—1260

Instruction Following Not comparable

Devstral Small 2505: —, Mercury: 65.2 (#224)

Instruction Following benchmarks
BenchmarkDevstral Small 2505Mercury
LMArena Instruction Following—1239

Long Context Not comparable

Devstral Small 2505: —, Mercury: 38.4 (#198)

Long Context benchmarks
BenchmarkDevstral Small 2505Mercury
LMArena Longer Query—1266

Writing & Preference Not comparable

Devstral Small 2505: —, Mercury: 46.2 (#221)

Writing & Preference benchmarks
BenchmarkDevstral Small 2505Mercury
LMArena Text—1282
LMArena Creative Writing—1191
LMArena Multi-Turn—1282

Frequently asked questions

Is Devstral Small 2505 better than Mercury?

Mercury is the stronger model overall, scoring 37.6 to 34.3 on the Noometry Index.

Is Devstral Small 2505 or Mercury better for coding?

They score almost the same on coding (38.9 vs 38.7); test both on your own repository before choosing.

How many benchmarks do Devstral Small 2505 and Mercury share?

1 benchmark has published results for both models. Devstral Small 2505 has 4 scored results on Noometry and Mercury has 9.

Related comparisons

Go deeper