Model comparison

Devstral Small 2505 vs Qwen2.5-VL 72B Instruct

Devstral Small 2505 is the stronger model overall, scoring 34.3 to 29.9 on the Noometry Index.

Last verified . 1 shared benchmarks.

Devstral Small 2505 Mistral AI

34.3

Rank #233 Reported

Summary

  • They share 1 benchmark with published results for both. Devstral Small 2505 scores higher in 0 categories and Qwen2.5-VL 72B Instruct in 1 category, but none of those gaps is larger than the uncertainty.
  • Devstral Small 2505 is cheaper at $0.10 / $0.30 per million input/output tokens, against $2.80 / $8.40 for Qwen2.5-VL 72B Instruct.
  • Qwen2.5-VL 72B Instruct accepts more context: 131K tokens versus 128K.

Side by side

Devstral Small 2505 and Qwen2.5-VL 72B Instruct specifications
Devstral Small 2505Qwen2.5-VL 72B Instruct
ProviderMistral AIAlibaba (Qwen)
Noometry Index34.329.9
Released2025-05-072024-09
WeightsOpenOpen
Context window128K131K
Max output128K8K
Input $ / M tokens$0.10$2.80
Output $ / M tokens$0.30$8.40
Results tracked46

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Devstral Small 2505: 38.9 (#166), Qwen2.5-VL 72B Instruct: —

Coding benchmarks
BenchmarkDevstral Small 2505Qwen2.5-VL 72B Instruct
SWE-bench Verified (bash only)56.4%—
SciCode28.8%—

Agentic & Tool Use Not comparable

Devstral Small 2505: —, Qwen2.5-VL 72B Instruct: 18.6 (#144)

Agentic & Tool Use benchmarks
BenchmarkDevstral Small 2505Qwen2.5-VL 72B Instruct
OSWorld—5%

Reasoning Too close to call

Devstral Small 2505: 19.7 (#252), Qwen2.5-VL 72B Instruct: 20.7 (#233)

Reasoning benchmarks
BenchmarkDevstral Small 2505Qwen2.5-VL 72B Instruct
Kagi LLM Benchmark37.7%36%
CritPt0%—

Multimodal Not comparable

Devstral Small 2505: —, Qwen2.5-VL 72B Instruct: 33.5 (#97)

Multimodal benchmarks
BenchmarkDevstral Small 2505Qwen2.5-VL 72B Instruct
LMArena Vision—1107
Video-MME—73.5%
GeoBench—62%
SpatialViz-Bench—33.3%

Frequently asked questions

Is Devstral Small 2505 better than Qwen2.5-VL 72B Instruct?

Devstral Small 2505 is the stronger model overall, scoring 34.3 to 29.9 on the Noometry Index.

Which is cheaper, Devstral Small 2505 or Qwen2.5-VL 72B Instruct?

Devstral Small 2505 is cheaper. It lists at $0.10 per million input tokens and $0.30 per million output tokens; Qwen2.5-VL 72B Instruct lists at $2.80 and $8.40.

Which has the bigger context window?

Qwen2.5-VL 72B Instruct does, with 131K tokens against 128K.

How many benchmarks do Devstral Small 2505 and Qwen2.5-VL 72B Instruct share?

1 benchmark has published results for both models. Devstral Small 2505 has 4 scored results on Noometry and Qwen2.5-VL 72B Instruct has 6.

Related comparisons

Go deeper