Model comparison

GLM-4.5-Air vs Qwen2.5 Plus 1127

GLM-4.5-Air and Qwen2.5 Plus 1127 score almost the same on the Noometry Index (38.9 vs 38.8), so choose on price, context window or the category you care about most.

Last verified . 14 shared benchmarks.

GLM-4.5-Air Z.ai (Zhipu)

38.9

Rank #177 Confirmed

Qwen2.5 Plus 1127 Alibaba (Qwen)

38.8

Rank #181 Confirmed

Summary

  • They share 14 benchmarks with published results for both. GLM-4.5-Air scores higher in 5 categories and Qwen2.5 Plus 1127 in 3 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where GLM-4.5-Air leads 49.1 to 41.9.
  • GLM-4.5-Air has downloadable open weights; the other is API-only.

Side by side

GLM-4.5-Air and Qwen2.5 Plus 1127 specifications
GLM-4.5-AirQwen2.5 Plus 1127
ProviderZ.ai (Zhipu)Alibaba (Qwen)
Noometry Index38.938.8
Released2025-07-20—
WeightsOpenProprietary
Context window131K—
Max output98K—
Input $ / M tokens$0.20—
Output $ / M tokens$1.10—
Results tracked2714

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 Plus 1127 leads

GLM-4.5-Air: 33.3 (#259), Qwen2.5 Plus 1127: 38.5 (#175)

Coding benchmarks
BenchmarkGLM-4.5-AirQwen2.5 Plus 1127
LMArena Coding13971314
GSO2.9%—

Reasoning Qwen2.5 Plus 1127 leads

GLM-4.5-Air: 24.1 (#166), Qwen2.5 Plus 1127: 25.9 (#141)

Reasoning benchmarks
BenchmarkGLM-4.5-AirQwen2.5 Plus 1127
LMArena Hard Prompts13791299
Kagi LLM Benchmark43%—
ForecastBench59.2—

Math Too close to call

GLM-4.5-Air: 36.2 (#170), Qwen2.5 Plus 1127: 36.1 (#174)

Math benchmarks
BenchmarkGLM-4.5-AirQwen2.5 Plus 1127
LMArena Math13961298
Omni-MATH39.1%—

Knowledge Too close to call

GLM-4.5-Air: 35.0 (#191), Qwen2.5 Plus 1127: 35.5 (#183)

Knowledge benchmarks
BenchmarkGLM-4.5-AirQwen2.5 Plus 1127
LMArena Expert13701289
Humanity's Last Exam8.1%—
MMLU-Pro76.2%—
Vectara Hallucination Rate9.3%—
GPQA (HELM)59.4%—

Multilingual GLM-4.5-Air leads

GLM-4.5-Air: 49.1 (#135), Qwen2.5 Plus 1127: 41.9 (#201)

Multilingual benchmarks
BenchmarkGLM-4.5-AirQwen2.5 Plus 1127
LMArena Non-English13661265
LMArena Chinese14261314
LMArena German13771231
LMArena Japanese13481207
LMArena Russian13731271
LMArena French1399—
LMArena Korean1308—
LMArena Spanish1386—

Instruction Following GLM-4.5-Air leads

GLM-4.5-Air: 69.6 (#171), Qwen2.5 Plus 1127: 67.2 (#199)

Instruction Following benchmarks
BenchmarkGLM-4.5-AirQwen2.5 Plus 1127
LMArena Instruction Following13541275
IFEval81.2%—

Long Context GLM-4.5-Air leads

GLM-4.5-Air: 41.6 (#135), Qwen2.5 Plus 1127: 39.2 (#184)

Long Context benchmarks
BenchmarkGLM-4.5-AirQwen2.5 Plus 1127
LMArena Longer Query13661292

Writing & Preference GLM-4.5-Air leads

GLM-4.5-Air: 55.9 (#139), Qwen2.5 Plus 1127: 49.4 (#192)

Writing & Preference benchmarks
BenchmarkGLM-4.5-AirQwen2.5 Plus 1127
LMArena Text13841299
LMArena Creative Writing13431262
LMArena Multi-Turn13711299
WildBench78.9%—

Frequently asked questions

Is GLM-4.5-Air better than Qwen2.5 Plus 1127?

GLM-4.5-Air and Qwen2.5 Plus 1127 score almost the same on the Noometry Index (38.9 vs 38.8), so choose on price, context window or the category you care about most.

Is GLM-4.5-Air or Qwen2.5 Plus 1127 better for coding?

Qwen2.5 Plus 1127 scores higher on coding benchmarks: 38.5 versus 33.3 in the Noometry coding category.

How many benchmarks do GLM-4.5-Air and Qwen2.5 Plus 1127 share?

14 benchmarks have published results for both models. GLM-4.5-Air has 27 scored results on Noometry and Qwen2.5 Plus 1127 has 14.

Related comparisons

Go deeper