Model comparison

Devstral Small 2505 vs GLM-4.5V

GLM-4.5V is the stronger model overall, scoring 39.8 to 34.3 on the Noometry Index. Devstral Small 2505 costs 6.0× less per token, which makes it the better buy when GLM-4.5V's lead doesn't matter for your workload.

Last verified . 1 shared benchmarks.

Devstral Small 2505 Mistral AI

34.3

Rank #233 Reported

GLM-4.5V Z.ai (Zhipu)

39.8

Rank #158 Confirmed

Summary

  • They share 1 benchmark with published results for both. Devstral Small 2505 scores higher in 0 categories and GLM-4.5V in 2 categories; one gap is clear of the uncertainty.
  • The widest gap is in reasoning, where GLM-4.5V leads 27.4 to 19.7.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 37.7% for Devstral Small 2505 and 59.8% for GLM-4.5V.
  • Devstral Small 2505 is cheaper at $0.10 / $0.30 per million input/output tokens, against $0.60 / $1.80 for GLM-4.5V.
  • Devstral Small 2505 accepts more context: 128K tokens versus 64K.

Side by side

Devstral Small 2505 and GLM-4.5V specifications
Devstral Small 2505GLM-4.5V
ProviderMistral AIZ.ai (Zhipu)
Noometry Index34.339.8
Released2025-05-072025-08-11
WeightsOpenOpen
Context window128K64K
Max output128K16K
Input $ / M tokens$0.10$0.60
Output $ / M tokens$0.30$1.80
Results tracked415

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Devstral Small 2505: 38.9 (#166), GLM-4.5V: 39.5 (#155)

Coding benchmarks
BenchmarkDevstral Small 2505GLM-4.5V
SWE-bench Verified (bash only)56.4%—
SciCode28.8%—
LMArena Coding—1347

Reasoning GLM-4.5V leads

Devstral Small 2505: 19.7 (#252), GLM-4.5V: 27.4 (#119)

Reasoning benchmarks
BenchmarkDevstral Small 2505GLM-4.5V
Kagi LLM Benchmark37.7%59.8%
CritPt0%—
LMArena Hard Prompts—1334

Math Not comparable

Devstral Small 2505: —, GLM-4.5V: 37.4 (#159)

Math benchmarks
BenchmarkDevstral Small 2505GLM-4.5V
LMArena Math—1354

Knowledge Not comparable

Devstral Small 2505: —, GLM-4.5V: 37.5 (#156)

Knowledge benchmarks
BenchmarkDevstral Small 2505GLM-4.5V
LMArena Expert—1353

Multimodal Not comparable

Devstral Small 2505: —, GLM-4.5V: 34.3 (#92)

Multimodal benchmarks
BenchmarkDevstral Small 2505GLM-4.5V
LMArena Vision—1154

Multilingual Not comparable

Devstral Small 2505: —, GLM-4.5V: 44.6 (#177)

Multilingual benchmarks
BenchmarkDevstral Small 2505GLM-4.5V
LMArena Non-English—1303
LMArena Chinese—1337
LMArena Russian—1298
LMArena Spanish—1336

Instruction Following Not comparable

Devstral Small 2505: —, GLM-4.5V: 69.2 (#175)

Instruction Following benchmarks
BenchmarkDevstral Small 2505GLM-4.5V
LMArena Instruction Following—1311

Long Context Not comparable

Devstral Small 2505: —, GLM-4.5V: 39.6 (#171)

Long Context benchmarks
BenchmarkDevstral Small 2505GLM-4.5V
LMArena Longer Query—1304

Writing & Preference Not comparable

Devstral Small 2505: —, GLM-4.5V: 52.5 (#170)

Writing & Preference benchmarks
BenchmarkDevstral Small 2505GLM-4.5V
LMArena Text—1333
LMArena Creative Writing—1295
LMArena Multi-Turn—1332

Frequently asked questions

Is Devstral Small 2505 better than GLM-4.5V?

GLM-4.5V is the stronger model overall, scoring 39.8 to 34.3 on the Noometry Index. Devstral Small 2505 costs 6.0× less per token, which makes it the better buy when GLM-4.5V's lead doesn't matter for your workload.

Which is cheaper, Devstral Small 2505 or GLM-4.5V?

Devstral Small 2505 is cheaper. It lists at $0.10 per million input tokens and $0.30 per million output tokens; GLM-4.5V lists at $0.60 and $1.80.

Is Devstral Small 2505 or GLM-4.5V better for coding?

They score almost the same on coding (38.9 vs 39.5); test both on your own repository before choosing.

Which has the bigger context window?

Devstral Small 2505 does, with 128K tokens against 64K.

How many benchmarks do Devstral Small 2505 and GLM-4.5V share?

1 benchmark has published results for both models. Devstral Small 2505 has 4 scored results on Noometry and GLM-4.5V has 15.

Related comparisons

Go deeper