Model comparison

Devstral Small 2505 vs Qwen3.5 Max Preview

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 34.3 on the Noometry Index.

Last verified . 0 shared benchmarks.

Devstral Small 2505 Mistral AI

34.3

Rank #233 Reported

Qwen3.5 Max Preview Alibaba (Qwen)

45.3

Rank #71 Confirmed

Summary

  • The widest gap is in reasoning, where Qwen3.5 Max Preview leads 30.8 to 19.7.
  • Devstral Small 2505 has downloadable open weights; the other is API-only.

Side by side

Devstral Small 2505 and Qwen3.5 Max Preview specifications
Devstral Small 2505Qwen3.5 Max Preview
ProviderMistral AIAlibaba (Qwen)
Noometry Index34.345.3
Released2025-05-07—
WeightsOpenProprietary
Context window128K—
Max output128K—
Input $ / M tokens$0.10—
Output $ / M tokens$0.30—
Results tracked417

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.5 Max Preview leads

Devstral Small 2505: 38.9 (#166), Qwen3.5 Max Preview: 44.0 (#77)

Coding benchmarks
BenchmarkDevstral Small 2505Qwen3.5 Max Preview
SWE-bench Verified (bash only)56.4%—
SciCode28.8%—
LMArena Coding—1487

Reasoning Qwen3.5 Max Preview leads

Devstral Small 2505: 19.7 (#252), Qwen3.5 Max Preview: 30.8 (#84)

Reasoning benchmarks
BenchmarkDevstral Small 2505Qwen3.5 Max Preview
Kagi LLM Benchmark37.7%—
CritPt0%—
LMArena Hard Prompts—1483

Math Not comparable

Devstral Small 2505: —, Qwen3.5 Max Preview: 40.1 (#94)

Math benchmarks
BenchmarkDevstral Small 2505Qwen3.5 Max Preview
LMArena Math—1474

Knowledge Not comparable

Devstral Small 2505: —, Qwen3.5 Max Preview: 41.8 (#107)

Knowledge benchmarks
BenchmarkDevstral Small 2505Qwen3.5 Max Preview
LMArena Expert—1489

Multilingual Not comparable

Devstral Small 2505: —, Qwen3.5 Max Preview: 56.2 (#22)

Multilingual benchmarks
BenchmarkDevstral Small 2505Qwen3.5 Max Preview
LMArena Non-English—1465
LMArena Chinese—1534
LMArena French—1484
LMArena German—1487
LMArena Japanese—1495
LMArena Korean—1438
LMArena Russian—1471
LMArena Spanish—1470

Instruction Following Not comparable

Devstral Small 2505: —, Qwen3.5 Max Preview: 77.0 (#31)

Instruction Following benchmarks
BenchmarkDevstral Small 2505Qwen3.5 Max Preview
LMArena Instruction Following—1467

Long Context Not comparable

Devstral Small 2505: —, Qwen3.5 Max Preview: 45.2 (#45)

Long Context benchmarks
BenchmarkDevstral Small 2505Qwen3.5 Max Preview
LMArena Longer Query—1476

Writing & Preference Not comparable

Devstral Small 2505: —, Qwen3.5 Max Preview: 66.0 (#41)

Writing & Preference benchmarks
BenchmarkDevstral Small 2505Qwen3.5 Max Preview
LMArena Text—1470
LMArena Creative Writing—1464
LMArena Multi-Turn—1478

Frequently asked questions

Is Devstral Small 2505 better than Qwen3.5 Max Preview?

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 34.3 on the Noometry Index.

Is Devstral Small 2505 or Qwen3.5 Max Preview better for coding?

Qwen3.5 Max Preview scores higher on coding benchmarks: 44.0 versus 38.9 in the Noometry coding category.

How many benchmarks do Devstral Small 2505 and Qwen3.5 Max Preview share?

0 benchmarks have published results for both models. Devstral Small 2505 has 4 scored results on Noometry and Qwen3.5 Max Preview has 17.

Related comparisons

Go deeper