Model comparison

GLM-5.1 vs Qwen3.7 Plus

GLM-5.1 is the stronger model overall, scoring 47.8 to 45.3 on the Noometry Index. Qwen3.7 Plus costs 3.1× less per token, which makes it the better buy when GLM-5.1's lead doesn't matter for your workload.

Last verified . 25 shared benchmarks.

GLM-5.1 Z.ai (Zhipu)

47.8

Rank #59 Confirmed

Qwen3.7 Plus Alibaba (Qwen)

45.3

Rank #72 Confirmed

Summary

  • They share 25 benchmarks with published results for both. GLM-5.1 scores higher in 6 categories and Qwen3.7 Plus in 3 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in coding, where GLM-5.1 leads 48.7 to 36.6.
  • The biggest single-benchmark swing is Chess Puzzles: 19% for GLM-5.1 and 24% for Qwen3.7 Plus.
  • Qwen3.7 Plus is cheaper at $0.40 / $1.60 per million input/output tokens, against $1.40 / $4.40 for GLM-5.1.
  • Qwen3.7 Plus accepts more context: 1M tokens versus 200K.
  • GLM-5.1 has downloadable open weights; the other is API-only.

Side by side

GLM-5.1 and Qwen3.7 Plus specifications
GLM-5.1Qwen3.7 Plus
ProviderZ.ai (Zhipu)Alibaba (Qwen)
Noometry Index47.845.3
Released2026-04-072026-06-02
WeightsOpenProprietary
Context window200K1M
Max output131K131K
Input $ / M tokens$1.40$0.40
Output $ / M tokens$4.40$1.60
Results tracked4132

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5.1 leads

GLM-5.1: 48.7 (#55), Qwen3.7 Plus: 36.6 (#206)

Coding benchmarks
BenchmarkGLM-5.1Qwen3.7 Plus
SciCode43.8%45.5%
LMArena Coding14851473
SWE-bench Verified74.2%—
FrontierCode—10.2%
LMArena WebDev1508—
WeirdML57.1%—
ALE-Bench887.1—

Agentic & Tool Use GLM-5.1 leads

GLM-5.1: 24.9 (#113), Qwen3.7 Plus: 21.4 (#138)

Agentic & Tool Use benchmarks
BenchmarkGLM-5.1Qwen3.7 Plus
APEX-Agents40.9%—
OSWorld 2.0—2.8%
ExploitBench18.1%—
GBAEval0%—
Vending-Bench 25,634—

Reasoning Too close to call

GLM-5.1: 39.1 (#60), Qwen3.7 Plus: 39.3 (#59)

Reasoning benchmarks
BenchmarkGLM-5.1Qwen3.7 Plus
NYT Connections (extended)77.7%74.8%
CritPt4.6%9.1%
Chess Puzzles19%24%
LMArena Hard Prompts14721460
Epoch Capabilities Index149.84147.37
SimpleBench55.1%—
Thematic Generalization69.8%—
Mystery Game Puzzles—17%
DTBench—84%
LMCA—37.6%

Math Too close to call

GLM-5.1: 49.7 (#60), Qwen3.7 Plus: 50.5 (#56)

Knowledge Too close to call

GLM-5.1: 54.9 (#50), Qwen3.7 Plus: 54.9 (#51)

Knowledge benchmarks
BenchmarkGLM-5.1Qwen3.7 Plus
GPQA Diamond89.9%87.9%
LMArena Expert14761467
SimpleQA Verified34%—

Multimodal Not comparable

GLM-5.1: —, Qwen3.7 Plus: 41.8 (#33)

Multimodal benchmarks
BenchmarkGLM-5.1Qwen3.7 Plus
LMArena Vision—1279
LMArena Document—1444

Multilingual Too close to call

GLM-5.1: 55.0 (#36), Qwen3.7 Plus: 54.8 (#38)

Multilingual benchmarks
BenchmarkGLM-5.1Qwen3.7 Plus
LMArena Non-English14471445
LMArena Chinese15151510
LMArena French14741473
LMArena German14651471
LMArena Japanese14341413
LMArena Korean14181415
LMArena Russian14541457
LMArena Spanish14691457

Instruction Following Too close to call

GLM-5.1: 76.3 (#42), Qwen3.7 Plus: 75.8 (#52)

Instruction Following benchmarks
BenchmarkGLM-5.1Qwen3.7 Plus
LMArena Instruction Following14511440

Long Context Too close to call

GLM-5.1: 44.9 (#53), Qwen3.7 Plus: 44.5 (#65)

Long Context benchmarks
BenchmarkGLM-5.1Qwen3.7 Plus
LMArena Longer Query14661455

Writing & Preference GLM-5.1 leads

GLM-5.1: 66.9 (#31), Qwen3.7 Plus: 64.3 (#56)

Writing & Preference benchmarks
BenchmarkGLM-5.1Qwen3.7 Plus
LMArena Text14611455
LMArena Creative Writing14531439
LMArena Multi-Turn14721460
EQ-Bench Creative Writing1592—

Frequently asked questions

Is GLM-5.1 better than Qwen3.7 Plus?

GLM-5.1 is the stronger model overall, scoring 47.8 to 45.3 on the Noometry Index. Qwen3.7 Plus costs 3.1× less per token, which makes it the better buy when GLM-5.1's lead doesn't matter for your workload.

Which is cheaper, GLM-5.1 or Qwen3.7 Plus?

Qwen3.7 Plus is cheaper. It lists at $0.40 per million input tokens and $1.60 per million output tokens; GLM-5.1 lists at $1.40 and $4.40.

Is GLM-5.1 or Qwen3.7 Plus better for coding?

GLM-5.1 scores higher on coding benchmarks: 48.7 versus 36.6 in the Noometry coding category.

Which has the bigger context window?

Qwen3.7 Plus does, with 1M tokens against 200K.

How many benchmarks do GLM-5.1 and Qwen3.7 Plus share?

25 benchmarks have published results for both models. GLM-5.1 has 41 scored results on Noometry and Qwen3.7 Plus has 32.

Related comparisons

Go deeper