Model comparison

Devstral Small 2505 vs Grok 4.1 Fast

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 34.3 on the Noometry Index. Devstral Small 2505 costs 1.8× less per token, which makes it the better buy when Grok 4.1 Fast's lead doesn't matter for your workload.

Last verified . 0 shared benchmarks.

Devstral Small 2505 Mistral AI

34.3

Rank #233 Reported

Grok 4.1 Fast xAI

41.4

Rank #136 Confirmed

Summary

  • The widest gap is in reasoning, where Grok 4.1 Fast leads 43.4 to 19.7.
  • Devstral Small 2505 is cheaper at $0.10 / $0.30 per million input/output tokens, against $0.20 / $0.50 for Grok 4.1 Fast.
  • Devstral Small 2505 has downloadable open weights; the other is API-only.

Side by side

Devstral Small 2505 and Grok 4.1 Fast specifications
Devstral Small 2505Grok 4.1 Fast
ProviderMistral AIxAI
Noometry Index34.341.4
Released2025-05-072025-06-27
WeightsOpenProprietary
Context window128K128K
Max output128K30K
Input $ / M tokens$0.10$0.20
Output $ / M tokens$0.30$0.50
Results tracked432

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Devstral Small 2505 leads

Devstral Small 2505: 38.9 (#166), Grok 4.1 Fast: 34.1 (#245)

Coding benchmarks
BenchmarkDevstral Small 2505Grok 4.1 Fast
SWE-bench Verified (bash only)56.4%—
LMArena WebDev—1242
SciCode28.8%—
LMArena Coding—1411
ALE-Bench—394.93

Agentic & Tool Use Not comparable

Devstral Small 2505: —, Grok 4.1 Fast: 36.3 (#39)

Agentic & Tool Use benchmarks
BenchmarkDevstral Small 2505Grok 4.1 Fast
Berkeley Function Calling Leaderboard—69.6%
τ²-bench Banking—13.1%
LMArena Search—1171
Vending-Bench 2—1,107

Reasoning Grok 4.1 Fast leads

Devstral Small 2505: 19.7 (#252), Grok 4.1 Fast: 43.4 (#49)

Reasoning benchmarks
BenchmarkDevstral Small 2505Grok 4.1 Fast
SimpleBench—56%
Kagi LLM Benchmark37.7%—
NYT Connections (extended)—87.4%
CritPt0%—
LMArena Hard Prompts—1407
DTBench—87.7%
ForecastBench—61

Math Not comparable

Devstral Small 2505: —, Grok 4.1 Fast: 31.9 (#221)

Math benchmarks
BenchmarkDevstral Small 2505Grok 4.1 Fast
MathArena Final-Answer Competitions—60.9%
ProofBench—4%
LMArena Math—1408

Knowledge Not comparable

Devstral Small 2505: —, Grok 4.1 Fast: 33.1 (#207)

Knowledge benchmarks
BenchmarkDevstral Small 2505Grok 4.1 Fast
Vectara Hallucination Rate—17.8%
LMArena Expert—1399

Multimodal Not comparable

Devstral Small 2505: —, Grok 4.1 Fast: 37.0 (#76)

Multimodal benchmarks
BenchmarkDevstral Small 2505Grok 4.1 Fast
LMArena Vision—1201

Multilingual Not comparable

Devstral Small 2505: —, Grok 4.1 Fast: 51.0 (#114)

Multilingual benchmarks
BenchmarkDevstral Small 2505Grok 4.1 Fast
LMArena Non-English—1391
LMArena Chinese—1441
LMArena French—1415
LMArena German—1404
LMArena Japanese—1349
LMArena Korean—1361
LMArena Russian—1387
LMArena Spanish—1413

Instruction Following Not comparable

Devstral Small 2505: —, Grok 4.1 Fast: 72.7 (#133)

Instruction Following benchmarks
BenchmarkDevstral Small 2505Grok 4.1 Fast
LMArena Instruction Following—1376

Long Context Not comparable

Devstral Small 2505: —, Grok 4.1 Fast: 42.4 (#126)

Long Context benchmarks
BenchmarkDevstral Small 2505Grok 4.1 Fast
LMArena Longer Query—1390

Writing & Preference Not comparable

Devstral Small 2505: —, Grok 4.1 Fast: 57.2 (#131)

Writing & Preference benchmarks
BenchmarkDevstral Small 2505Grok 4.1 Fast
LMArena Text—1408
LMArena Creative Writing—1394
EQ-Bench Creative Writing—1327
LMArena Multi-Turn—1389

Frequently asked questions

Is Devstral Small 2505 better than Grok 4.1 Fast?

Grok 4.1 Fast is the stronger model overall, scoring 41.4 to 34.3 on the Noometry Index. Devstral Small 2505 costs 1.8× less per token, which makes it the better buy when Grok 4.1 Fast's lead doesn't matter for your workload.

Which is cheaper, Devstral Small 2505 or Grok 4.1 Fast?

Devstral Small 2505 is cheaper. It lists at $0.10 per million input tokens and $0.30 per million output tokens; Grok 4.1 Fast lists at $0.20 and $0.50.

Is Devstral Small 2505 or Grok 4.1 Fast better for coding?

Devstral Small 2505 scores higher on coding benchmarks: 38.9 versus 34.1 in the Noometry coding category.

Which has the bigger context window?

Both accept 128K tokens.

How many benchmarks do Devstral Small 2505 and Grok 4.1 Fast share?

0 benchmarks have published results for both models. Devstral Small 2505 has 4 scored results on Noometry and Grok 4.1 Fast has 32.

Related comparisons

Go deeper