Model comparison

Devstral Small 2505 vs Qwen1.5-110B

Devstral Small 2505 and Qwen1.5-110B score almost the same on the Noometry Index (34.3 vs 34.2), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

Devstral Small 2505 Mistral AI

34.3

Rank #233 Reported

Qwen1.5-110B Alibaba (Qwen)

34.2

Rank #234 Confirmed

Summary

  • The widest gap is in coding, where Devstral Small 2505 leads 38.9 to 33.0.

Side by side

Devstral Small 2505 and Qwen1.5-110B specifications
Devstral Small 2505Qwen1.5-110B
ProviderMistral AIAlibaba (Qwen)
Noometry Index34.334.2
Released2025-05-072024-04-25
WeightsOpenOpen
Context window128K—
Max output128K—
Input $ / M tokens$0.10—
Output $ / M tokens$0.30—
Results tracked420

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Devstral Small 2505 leads

Devstral Small 2505: 38.9 (#166), Qwen1.5-110B: 33.0 (#264)

Coding benchmarks
BenchmarkDevstral Small 2505Qwen1.5-110B
SWE-bench Verified (bash only)56.4%—
SciCode28.8%—
BigCodeBench Instruct—35%
LMArena Coding—1184
BigCodeBench Complete—44.4%

Reasoning Qwen1.5-110B leads

Devstral Small 2505: 19.7 (#252), Qwen1.5-110B: 22.7 (#189)

Reasoning benchmarks
BenchmarkDevstral Small 2505Qwen1.5-110B
Kagi LLM Benchmark37.7%—
CritPt0%—
LMArena Hard Prompts—1168
ForecastBench—57.7

Math Not comparable

Devstral Small 2505: —, Qwen1.5-110B: 33.7 (#201)

Math benchmarks
BenchmarkDevstral Small 2505Qwen1.5-110B
LMArena Math—1185

Knowledge Not comparable

Devstral Small 2505: —, Qwen1.5-110B: 31.2 (#219)

Knowledge benchmarks
BenchmarkDevstral Small 2505Qwen1.5-110B
LMArena Expert—1144

Multilingual Not comparable

Devstral Small 2505: —, Qwen1.5-110B: 33.6 (#250)

Multilingual benchmarks
BenchmarkDevstral Small 2505Qwen1.5-110B
LMArena Non-English—1142
LMArena Chinese—1206
LMArena French—1151
LMArena German—1123
LMArena Japanese—1074
LMArena Korean—1044
LMArena Russian—1118
LMArena Spanish—1142

Instruction Following Not comparable

Devstral Small 2505: —, Qwen1.5-110B: 60.3 (#252)

Instruction Following benchmarks
BenchmarkDevstral Small 2505Qwen1.5-110B
LMArena Instruction Following—1158

Long Context Not comparable

Devstral Small 2505: —, Qwen1.5-110B: 35.1 (#242)

Long Context benchmarks
BenchmarkDevstral Small 2505Qwen1.5-110B
LMArena Longer Query—1157

Writing & Preference Not comparable

Devstral Small 2505: —, Qwen1.5-110B: 38.0 (#255)

Writing & Preference benchmarks
BenchmarkDevstral Small 2505Qwen1.5-110B
LMArena Text—1175
LMArena Creative Writing—1148
LMArena Multi-Turn—1160

Frequently asked questions

Is Devstral Small 2505 better than Qwen1.5-110B?

Devstral Small 2505 and Qwen1.5-110B score almost the same on the Noometry Index (34.3 vs 34.2), so choose on price, context window or the category you care about most.

Is Devstral Small 2505 or Qwen1.5-110B better for coding?

Devstral Small 2505 scores higher on coding benchmarks: 38.9 versus 33.0 in the Noometry coding category.

How many benchmarks do Devstral Small 2505 and Qwen1.5-110B share?

0 benchmarks have published results for both models. Devstral Small 2505 has 4 scored results on Noometry and Qwen1.5-110B has 20.

Related comparisons

Go deeper