Model comparison

Devstral Small 2505 vs GPT-4.1 mini

Devstral Small 2505 and GPT-4.1 mini score almost the same on the Noometry Index (34.3 vs 33.6), so choose on price, context window or the category you care about most.

Last verified . 4 shared benchmarks.

Devstral Small 2505 Mistral AI

34.3

Rank #233 Reported

GPT-4.1 mini OpenAI

33.6

Rank #240 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Devstral Small 2505 scores higher in 2 categories and GPT-4.1 mini in 0 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Devstral Small 2505 leads 19.7 to 10.8.
  • The biggest single-benchmark swing is SWE-bench Verified (bash only): 56.4% for Devstral Small 2505 and 23.9% for GPT-4.1 mini.
  • Devstral Small 2505 is cheaper at $0.10 / $0.30 per million input/output tokens, against $0.40 / $1.60 for GPT-4.1 mini.
  • GPT-4.1 mini accepts more context: 1.05M tokens versus 128K.
  • Devstral Small 2505 has downloadable open weights; the other is API-only.

Side by side

Devstral Small 2505 and GPT-4.1 mini specifications
Devstral Small 2505GPT-4.1 mini
ProviderMistral AIOpenAI
Noometry Index34.333.6
Released2025-05-072025-04-14
WeightsOpenProprietary
Context window128K1.05M
Max output128K33K
Input $ / M tokens$0.10$0.40
Output $ / M tokens$0.30$1.60
Results tracked447

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Devstral Small 2505 leads

Devstral Small 2505: 38.9 (#166), GPT-4.1 mini: 30.6 (#293)

Coding benchmarks
BenchmarkDevstral Small 2505GPT-4.1 mini
SWE-bench Verified (bash only)56.4%23.9%
SciCode28.8%40.4%
Aider Polyglot—32.4%
WeirdML—37.6%
BigCodeBench Instruct—48.9%
LMArena Coding—1367
CadEval—16%

Agentic & Tool Use Not comparable

Devstral Small 2505: —, GPT-4.1 mini: 33.3 (#55)

Agentic & Tool Use benchmarks
BenchmarkDevstral Small 2505GPT-4.1 mini
Berkeley Function Calling Leaderboard—50.5%

Reasoning Devstral Small 2505 leads

Devstral Small 2505: 19.7 (#252), GPT-4.1 mini: 10.8 (#340)

Reasoning benchmarks
BenchmarkDevstral Small 2505GPT-4.1 mini
Kagi LLM Benchmark37.7%48.6%
CritPt0%0%
ARC-AGI-2—0%
ARC-AGI-1—3.5%
Chess Puzzles—7%
LMArena Hard Prompts—1349
Mystery Game Puzzles—7%
DTBench—68.8%
LMCA—21.1%
Epoch Capabilities Index—135.01

Math Not comparable

Devstral Small 2505: —, GPT-4.1 mini: 24.1 (#270)

Math benchmarks
BenchmarkDevstral Small 2505GPT-4.1 mini
FrontierMath (Tiers 1-3)—6.7%
OTIS Mock AIME 2024-2025—44.7%
Omni-MATH—49.1%
LMArena Math—1343
MATH Level 5—87.3%
FrontierMath (Feb 2025 set)—4.5%

Knowledge Not comparable

Devstral Small 2505: —, GPT-4.1 mini: 34.7 (#194)

Knowledge benchmarks
BenchmarkDevstral Small 2505GPT-4.1 mini
GPQA Diamond—65.8%
SimpleQA Verified—12.7%
MMLU-Pro—78.3%
GPQA (HELM)—61.4%
LMArena Expert—1338

Multimodal Not comparable

Devstral Small 2505: —, GPT-4.1 mini: 35.8 (#82)

Multimodal benchmarks
BenchmarkDevstral Small 2505GPT-4.1 mini
LMArena Vision—1181

Multilingual Not comparable

Devstral Small 2505: —, GPT-4.1 mini: 45.7 (#166)

Multilingual benchmarks
BenchmarkDevstral Small 2505GPT-4.1 mini
LMArena Non-English—1318
LMArena Chinese—1329
LMArena French—1358
LMArena German—1351
LMArena Japanese—1290
LMArena Korean—1298
LMArena Russian—1324
LMArena Spanish—1319

Instruction Following Not comparable

Devstral Small 2505: —, GPT-4.1 mini: 73.7 (#118)

Instruction Following benchmarks
BenchmarkDevstral Small 2505GPT-4.1 mini
IFEval—90.4%
LMArena Instruction Following—1333

Long Context Not comparable

Devstral Small 2505: —, GPT-4.1 mini: 31.8 (#275)

Long Context benchmarks
BenchmarkDevstral Small 2505GPT-4.1 mini
Fiction.LiveBench—44.4%
LMArena Longer Query—1344

Writing & Preference Not comparable

Devstral Small 2505: —, GPT-4.1 mini: 48.6 (#199)

Writing & Preference benchmarks
BenchmarkDevstral Small 2505GPT-4.1 mini
LMArena Text—1340
LMArena Creative Writing—1300
EQ-Bench Creative Writing—1147
WildBench—83.8%
LMArena Multi-Turn—1354

Frequently asked questions

Is Devstral Small 2505 better than GPT-4.1 mini?

Devstral Small 2505 and GPT-4.1 mini score almost the same on the Noometry Index (34.3 vs 33.6), so choose on price, context window or the category you care about most.

Which is cheaper, Devstral Small 2505 or GPT-4.1 mini?

Devstral Small 2505 is cheaper. It lists at $0.10 per million input tokens and $0.30 per million output tokens; GPT-4.1 mini lists at $0.40 and $1.60.

Is Devstral Small 2505 or GPT-4.1 mini better for coding?

Devstral Small 2505 scores higher on coding benchmarks: 38.9 versus 30.6 in the Noometry coding category.

Which has the bigger context window?

GPT-4.1 mini does, with 1.05M tokens against 128K.

How many benchmarks do Devstral Small 2505 and GPT-4.1 mini share?

4 benchmarks have published results for both models. Devstral Small 2505 has 4 scored results on Noometry and GPT-4.1 mini has 47.

Related comparisons

Go deeper