Model comparison

Devstral Small 2505 vs Mercury 2

Mercury 2 is the stronger model overall, scoring 39.1 to 34.3 on the Noometry Index. Devstral Small 2505 costs 2.5× less per token, which makes it the better buy when Mercury 2's lead doesn't matter for your workload.

Last verified . 2 shared benchmarks.

Devstral Small 2505 Mistral AI

34.3

Rank #233 Reported

Mercury 2 Inception

39.1

Rank #175 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Devstral Small 2505 scores higher in 1 category and Mercury 2 in 1 category; 2 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Devstral Small 2505 leads 38.9 to 33.5.
  • The biggest single-benchmark swing is SciCode: 28.8% for Devstral Small 2505 and 38.7% for Mercury 2.
  • Devstral Small 2505 is cheaper at $0.10 / $0.30 per million input/output tokens, against $0.25 / $0.75 for Mercury 2.
  • Devstral Small 2505 has downloadable open weights; the other is API-only.

Side by side

Devstral Small 2505 and Mercury 2 specifications
Devstral Small 2505Mercury 2
ProviderMistral AIInception
Noometry Index34.339.1
Released2025-05-072026-02-20
WeightsOpenProprietary
Context window128K128K
Max output128K50K
Input $ / M tokens$0.10$0.25
Output $ / M tokens$0.30$0.75
Results tracked417

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Devstral Small 2505 leads

Devstral Small 2505: 38.9 (#166), Mercury 2: 33.5 (#255)

Coding benchmarks
BenchmarkDevstral Small 2505Mercury 2
SciCode28.8%38.7%
SWE-bench Verified (bash only)56.4%—
LMArena WebDev—1171
WeirdML—43.2%
LMArena Coding—1391
ALE-Bench—785.58

Reasoning Mercury 2 leads

Devstral Small 2505: 19.7 (#252), Mercury 2: 23.8 (#170)

Reasoning benchmarks
BenchmarkDevstral Small 2505Mercury 2
CritPt0%0.8%
Kagi LLM Benchmark37.7%—
LMArena Hard Prompts—1362

Knowledge Not comparable

Devstral Small 2505: —, Mercury 2: 36.2 (#172)

Knowledge benchmarks
BenchmarkDevstral Small 2505Mercury 2
Vectara Hallucination Rate—12.3%
LMArena Expert—1358

Multilingual Not comparable

Devstral Small 2505: —, Mercury 2: 46.6 (#157)

Multilingual benchmarks
BenchmarkDevstral Small 2505Mercury 2
LMArena Non-English—1331
LMArena Chinese—1417
LMArena Russian—1304

Instruction Following Not comparable

Devstral Small 2505: —, Mercury 2: 70.2 (#165)

Instruction Following benchmarks
BenchmarkDevstral Small 2505Mercury 2
LMArena Instruction Following—1329

Long Context Not comparable

Devstral Small 2505: —, Mercury 2: 40.5 (#154)

Long Context benchmarks
BenchmarkDevstral Small 2505Mercury 2
LMArena Longer Query—1330

Writing & Preference Not comparable

Devstral Small 2505: —, Mercury 2: 53.8 (#155)

Writing & Preference benchmarks
BenchmarkDevstral Small 2505Mercury 2
LMArena Text—1355
LMArena Creative Writing—1289
LMArena Multi-Turn—1358

Frequently asked questions

Is Devstral Small 2505 better than Mercury 2?

Mercury 2 is the stronger model overall, scoring 39.1 to 34.3 on the Noometry Index. Devstral Small 2505 costs 2.5× less per token, which makes it the better buy when Mercury 2's lead doesn't matter for your workload.

Which is cheaper, Devstral Small 2505 or Mercury 2?

Devstral Small 2505 is cheaper. It lists at $0.10 per million input tokens and $0.30 per million output tokens; Mercury 2 lists at $0.25 and $0.75.

Is Devstral Small 2505 or Mercury 2 better for coding?

Devstral Small 2505 scores higher on coding benchmarks: 38.9 versus 33.5 in the Noometry coding category.

Which has the bigger context window?

Both accept 128K tokens.

How many benchmarks do Devstral Small 2505 and Mercury 2 share?

2 benchmarks have published results for both models. Devstral Small 2505 has 4 scored results on Noometry and Mercury 2 has 17.

Related comparisons

Go deeper