Model comparison

Devstral Small 2505 vs GLM-4.7-Flash

GLM-4.7-Flash is the stronger model overall, scoring 38.8 to 34.3 on the Noometry Index.

Last verified . 0 shared benchmarks.

Devstral Small 2505 Mistral AI

34.3

Rank #233 Reported

GLM-4.7-Flash Z.ai (Zhipu)

38.8

Rank #180 Confirmed

Summary

  • Both cost about the same: $0.10 input and $0.30 output per million tokens.
  • GLM-4.7-Flash accepts more context: 200K tokens versus 128K.

Side by side

Devstral Small 2505 and GLM-4.7-Flash specifications
Devstral Small 2505GLM-4.7-Flash
ProviderMistral AIZ.ai (Zhipu)
Noometry Index34.338.8
Released2025-05-072026-01-19
WeightsOpenOpen
Context window128K200K
Max output128K131K
Input $ / M tokens$0.10$0.06
Output $ / M tokens$0.30$0.40
Results tracked421

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-4.7-Flash leads

Devstral Small 2505: 38.9 (#166), GLM-4.7-Flash: 40.6 (#135)

Coding benchmarks
BenchmarkDevstral Small 2505GLM-4.7-Flash
SWE-bench Verified (bash only)56.4%—
SciCode28.8%—
LMArena Coding—1383

Reasoning GLM-4.7-Flash leads

Devstral Small 2505: 19.7 (#252), GLM-4.7-Flash: 20.9 (#229)

Reasoning benchmarks
BenchmarkDevstral Small 2505GLM-4.7-Flash
Kagi LLM Benchmark37.7%—
CritPt0%—
Chess Puzzles—0%
LMArena Hard Prompts—1356

Math Not comparable

Devstral Small 2505: —, GLM-4.7-Flash: 36.1 (#173)

Math benchmarks
BenchmarkDevstral Small 2505GLM-4.7-Flash
OTIS Mock AIME 2024-2025—58.3%
LMArena Math—1355

Knowledge Not comparable

Devstral Small 2505: —, GLM-4.7-Flash: 35.5 (#184)

Knowledge benchmarks
BenchmarkDevstral Small 2505GLM-4.7-Flash
GPQA Diamond—60.5%
Vectara Hallucination Rate—9.3%
LMArena Expert—1357

Multilingual Not comparable

Devstral Small 2505: —, GLM-4.7-Flash: 46.5 (#158)

Multilingual benchmarks
BenchmarkDevstral Small 2505GLM-4.7-Flash
LMArena Non-English—1330
LMArena Chinese—1403
LMArena French—1332
LMArena German—1337
LMArena Korean—1283
LMArena Russian—1332
LMArena Spanish—1350

Instruction Following Not comparable

Devstral Small 2505: —, GLM-4.7-Flash: 70.1 (#167)

Instruction Following benchmarks
BenchmarkDevstral Small 2505GLM-4.7-Flash
LMArena Instruction Following—1327

Long Context Not comparable

Devstral Small 2505: —, GLM-4.7-Flash: 40.9 (#148)

Long Context benchmarks
BenchmarkDevstral Small 2505GLM-4.7-Flash
LMArena Longer Query—1345

Writing & Preference Not comparable

Devstral Small 2505: —, GLM-4.7-Flash: 47.4 (#210)

Writing & Preference benchmarks
BenchmarkDevstral Small 2505GLM-4.7-Flash
LMArena Text—1351
LMArena Creative Writing—1297
EQ-Bench Creative Writing—1125
LMArena Multi-Turn—1342

Frequently asked questions

Is Devstral Small 2505 better than GLM-4.7-Flash?

GLM-4.7-Flash is the stronger model overall, scoring 38.8 to 34.3 on the Noometry Index.

Which is cheaper, Devstral Small 2505 or GLM-4.7-Flash?

GLM-4.7-Flash is cheaper. It lists at $0.06 per million input tokens and $0.40 per million output tokens; Devstral Small 2505 lists at $0.10 and $0.30.

Is Devstral Small 2505 or GLM-4.7-Flash better for coding?

GLM-4.7-Flash scores higher on coding benchmarks: 40.6 versus 38.9 in the Noometry coding category.

Which has the bigger context window?

GLM-4.7-Flash does, with 200K tokens against 128K.

How many benchmarks do Devstral Small 2505 and GLM-4.7-Flash share?

0 benchmarks have published results for both models. Devstral Small 2505 has 4 scored results on Noometry and GLM-4.7-Flash has 21.

Related comparisons

Go deeper