Model comparison

Devstral Small 2505 vs GPT-4o mini

Devstral Small 2505 is the stronger model overall, scoring 34.3 to 25.5 on the Noometry Index.

Last verified . 1 shared benchmarks.

Devstral Small 2505 Mistral AI

34.3

Rank #233 Reported

GPT-4o mini OpenAI

25.5

Rank #343 Confirmed

Summary

  • They share 1 benchmark with published results for both. Devstral Small 2505 scores higher in 2 categories and GPT-4o mini in 0 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Devstral Small 2505 leads 38.9 to 22.0.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 37.7% for Devstral Small 2505 and 28.8% for GPT-4o mini.
  • Devstral Small 2505 is cheaper at $0.10 / $0.30 per million input/output tokens, against $0.15 / $0.60 for GPT-4o mini.
  • Devstral Small 2505 has downloadable open weights; the other is API-only.

Side by side

Devstral Small 2505 and GPT-4o mini specifications
Devstral Small 2505GPT-4o mini
ProviderMistral AIOpenAI
Noometry Index34.325.5
Released2025-05-072024-07-18
WeightsOpenProprietary
Context window128K128K
Max output128K16K
Input $ / M tokens$0.10$0.15
Output $ / M tokens$0.30$0.60
Results tracked460

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Devstral Small 2505 leads

Devstral Small 2505: 38.9 (#166), GPT-4o mini: 22.0 (#335)

Coding benchmarks
BenchmarkDevstral Small 2505GPT-4o mini
SWE-bench Verified (bash only)56.4%—
Aider Polyglot—3.6%
SciCode28.8%—
WeirdML—11.8%
BigCodeBench Instruct—46.1%
LiveBench Coding—43.1%
LMArena Coding—1290
BigCodeBench Complete—57.4%
HumanEval+—83.5%
MBPP+—72.2%

Agentic & Tool Use Not comparable

Devstral Small 2505: —, GPT-4o mini: 27.5 (#101)

Agentic & Tool Use benchmarks
BenchmarkDevstral Small 2505GPT-4o mini
BALROG—17.4%

Reasoning Devstral Small 2505 leads

Devstral Small 2505: 19.7 (#252), GPT-4o mini: 8.7 (#347)

Reasoning benchmarks
BenchmarkDevstral Small 2505GPT-4o mini
Kagi LLM Benchmark37.7%28.8%
ARC-AGI-2—0%
SimpleBench—10.7%
CritPt0%—
Chess Puzzles—0%
LiveBench Reasoning—32.8%
LMArena Hard Prompts—1267
Mystery Game Puzzles—12%
DTBench—54.4%
LiveBench Data Analysis—50%
LMCA—10.4%
Epoch Capabilities Index—126.56
LiveBench—41.3%
PIQA—88.7%

Math Not comparable

Devstral Small 2505: —, GPT-4o mini: 10.4 (#314)

Math benchmarks
BenchmarkDevstral Small 2505GPT-4o mini
FrontierMath (Tiers 1-3)—0.7%
OTIS Mock AIME 2024-2025—6.9%
Omni-MATH—28%
LiveBench Math—36.3%
LMArena Math—1267
MATH Level 5—52.6%
GSM8K—91.3%

Knowledge Not comparable

Devstral Small 2505: —, GPT-4o mini: 17.7 (#284)

Knowledge benchmarks
BenchmarkDevstral Small 2505GPT-4o mini
GPQA Diamond—37.7%
SimpleQA Verified—8.3%
MMLU-Pro—60.3%
Confabulations—37.2%
GPQA (HELM)—36.8%
LMArena Expert—1235
BoolQ—88.7%
MMLU—81.8%

Multimodal Not comparable

Devstral Small 2505: —, GPT-4o mini: 25.9 (#122)

Multimodal benchmarks
BenchmarkDevstral Small 2505GPT-4o mini
LMArena Vision—1066
Video-MME—64.8%
GeoBench—64%
VPCT—34%

Multilingual Not comparable

Devstral Small 2505: —, GPT-4o mini: 42.0 (#199)

Multilingual benchmarks
BenchmarkDevstral Small 2505GPT-4o mini
LMArena Non-English—1266
LMArena Chinese—1265
LMArena French—1297
LMArena German—1272
LMArena Japanese—1216
LMArena Korean—1195
LMArena Russian—1275
LMArena Spanish—1276

Instruction Following Not comparable

Devstral Small 2505: —, GPT-4o mini: 61.9 (#239)

Instruction Following benchmarks
BenchmarkDevstral Small 2505GPT-4o mini
LiveBench Instruction Following—56.8%
IFEval—78.2%
LMArena Instruction Following—1258

Long Context Not comparable

Devstral Small 2505: —, GPT-4o mini: 39.1 (#186)

Long Context benchmarks
BenchmarkDevstral Small 2505GPT-4o mini
LMArena Longer Query—1289

Writing & Preference Not comparable

Devstral Small 2505: —, GPT-4o mini: 39.5 (#248)

Writing & Preference benchmarks
BenchmarkDevstral Small 2505GPT-4o mini
LMArena Text—1286
LMArena Creative Writing—1268
Short-Story Creative Writing—67.2%
EQ-Bench Creative Writing—873
WildBench—79.1%
LMArena Multi-Turn—1285
LiveBench Language—28.6%

Frequently asked questions

Is Devstral Small 2505 better than GPT-4o mini?

Devstral Small 2505 is the stronger model overall, scoring 34.3 to 25.5 on the Noometry Index.

Which is cheaper, Devstral Small 2505 or GPT-4o mini?

Devstral Small 2505 is cheaper. It lists at $0.10 per million input tokens and $0.30 per million output tokens; GPT-4o mini lists at $0.15 and $0.60.

Is Devstral Small 2505 or GPT-4o mini better for coding?

Devstral Small 2505 scores higher on coding benchmarks: 38.9 versus 22.0 in the Noometry coding category.

Which has the bigger context window?

Both accept 128K tokens.

How many benchmarks do Devstral Small 2505 and GPT-4o mini share?

1 benchmark has published results for both models. Devstral Small 2505 has 4 scored results on Noometry and GPT-4o mini has 60.

Related comparisons

Go deeper