Model comparison
Devstral Small 2505 vs Grok Build 0.1
Grok Build 0.1 is the stronger model overall, scoring 36.4 to 34.3 on the Noometry Index. Devstral Small 2505 costs 8.3× less per token, which makes it the better buy when Grok Build 0.1's lead doesn't matter for your workload.
Last verified . 2 shared benchmarks.
Summary
- They share 2 benchmarks with published results for both. Devstral Small 2505 scores higher in 0 categories and Grok Build 0.1 in 2 categories; 2 gaps are clear of the uncertainty.
- The widest gap is in reasoning, where Grok Build 0.1 leads 32.2 to 19.7.
- The biggest single-benchmark swing is SciCode: 28.8% for Devstral Small 2505 and 50.2% for Grok Build 0.1.
- Devstral Small 2505 is cheaper at $0.10 / $0.30 per million input/output tokens, against $1 / $2 for Grok Build 0.1.
- Grok Build 0.1 accepts more context: 256K tokens versus 128K.
- Devstral Small 2505 has downloadable open weights; the other is API-only.
Side by side
| Devstral Small 2505 | Grok Build 0.1 | |
|---|---|---|
| Provider | Mistral AI | xAI |
| Noometry Index | 34.3 | 36.4 |
| Released | 2025-05-07 | 2026-04-16 |
| Weights | Open | Proprietary |
| Context window | 128K | 256K |
| Max output | 128K | 256K |
| Input $ / M tokens | $0.10 | $1 |
| Output $ / M tokens | $0.30 | $2 |
| Results tracked | 4 | 3 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Grok Build 0.1 leads
Devstral Small 2505: 38.9 (#166), Grok Build 0.1: 43.1 (#91)
| Benchmark | Devstral Small 2505 | Grok Build 0.1 |
|---|---|---|
| SciCode | 28.8% | 50.2% |
| SWE-bench Verified (bash only) | 56.4% | — |
Agentic & Tool Use Not comparable
Devstral Small 2505: —, Grok Build 0.1: 22.7 (#129)
| Benchmark | Devstral Small 2505 | Grok Build 0.1 |
|---|---|---|
| GBAEval | — | 2.4% |
Reasoning Grok Build 0.1 leads
Devstral Small 2505: 19.7 (#252), Grok Build 0.1: 32.2 (#77)
| Benchmark | Devstral Small 2505 | Grok Build 0.1 |
|---|---|---|
| CritPt | 0% | 9.1% |
| Kagi LLM Benchmark | 37.7% | — |
Frequently asked questions
Is Devstral Small 2505 better than Grok Build 0.1?
Grok Build 0.1 is the stronger model overall, scoring 36.4 to 34.3 on the Noometry Index. Devstral Small 2505 costs 8.3× less per token, which makes it the better buy when Grok Build 0.1's lead doesn't matter for your workload.
Which is cheaper, Devstral Small 2505 or Grok Build 0.1?
Devstral Small 2505 is cheaper. It lists at $0.10 per million input tokens and $0.30 per million output tokens; Grok Build 0.1 lists at $1 and $2.
Is Devstral Small 2505 or Grok Build 0.1 better for coding?
Grok Build 0.1 scores higher on coding benchmarks: 43.1 versus 38.9 in the Noometry coding category.
Which has the bigger context window?
Grok Build 0.1 does, with 256K tokens against 128K.
How many benchmarks do Devstral Small 2505 and Grok Build 0.1 share?
2 benchmarks have published results for both models. Devstral Small 2505 has 4 scored results on Noometry and Grok Build 0.1 has 3.