Model comparison
Devstral Small 2505 vs Mercury 2.5
Devstral Small 2505 and Mercury 2.5 score almost the same on the Noometry Index (34.3 vs 33.5), so choose on price, context window or the category you care about most.
Last verified . 2 shared benchmarks.
Summary
- They share 2 benchmarks with published results for both. Devstral Small 2505 scores higher in 0 categories and Mercury 2.5 in 2 categories; one gap is clear of the uncertainty.
- The biggest single-benchmark swing is SciCode: 28.8% for Devstral Small 2505 and 38.5% for Mercury 2.5.
- Mercury 2.5 is cheaper at $0.04 / $0.15 per million input/output tokens, against $0.10 / $0.30 for Devstral Small 2505.
- Mercury 2.5 accepts more context: 260K tokens versus 128K.
- Devstral Small 2505 has downloadable open weights; the other is API-only.
Side by side
| Devstral Small 2505 | Mercury 2.5 | |
|---|---|---|
| Provider | Mistral AI | Inception |
| Noometry Index | 34.3 | 33.5 |
| Released | 2025-05-07 | 2026-09-08 |
| Weights | Open | Proprietary |
| Context window | 128K | 260K |
| Max output | 128K | 66K |
| Input $ / M tokens | $0.10 | $0.04 |
| Output $ / M tokens | $0.30 | $0.15 |
| Results tracked | 4 | 4 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Too close to call
Devstral Small 2505: 38.9 (#166), Mercury 2.5: 39.5 (#156)
| Benchmark | Devstral Small 2505 | Mercury 2.5 |
|---|---|---|
| SciCode | 28.8% | 38.5% |
| SWE-bench Verified (bash only) | 56.4% | — |
| ALE-Bench | — | 301.65 |
Reasoning Mercury 2.5 leads
Devstral Small 2505: 19.7 (#252), Mercury 2.5: 22.4 (#193)
| Benchmark | Devstral Small 2505 | Mercury 2.5 |
|---|---|---|
| CritPt | 0% | 0% |
| Kagi LLM Benchmark | 37.7% | — |
Math Not comparable
Devstral Small 2505: —, Mercury 2.5: 23.3 (#272)
| Benchmark | Devstral Small 2505 | Mercury 2.5 |
|---|---|---|
| ProofBench | — | 3% |
Frequently asked questions
Is Devstral Small 2505 better than Mercury 2.5?
Devstral Small 2505 and Mercury 2.5 score almost the same on the Noometry Index (34.3 vs 33.5), so choose on price, context window or the category you care about most.
Which is cheaper, Devstral Small 2505 or Mercury 2.5?
Mercury 2.5 is cheaper. It lists at $0.04 per million input tokens and $0.15 per million output tokens; Devstral Small 2505 lists at $0.10 and $0.30.
Is Devstral Small 2505 or Mercury 2.5 better for coding?
They score almost the same on coding (38.9 vs 39.5); test both on your own repository before choosing.
Which has the bigger context window?
Mercury 2.5 does, with 260K tokens against 128K.
How many benchmarks do Devstral Small 2505 and Mercury 2.5 share?
2 benchmarks have published results for both models. Devstral Small 2505 has 4 scored results on Noometry and Mercury 2.5 has 4.