Model comparison

GLM-5.2 vs Qwen3.6 Max Preview

GLM-5.2 and Qwen3.6 Max Preview score almost the same on the Noometry Index (51.1 vs 51.5), so choose on price, context window or the category you care about most.

Last verified . 27 shared benchmarks.

GLM-5.2 Z.ai (Zhipu)

51.1

Rank #44 Confirmed

Qwen3.6 Max Preview Alibaba (Qwen)

51.5

Rank #43 Confirmed

Summary

  • They share 27 benchmarks with published results for both. GLM-5.2 scores higher in 7 categories and Qwen3.6 Max Preview in 1 category; 5 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where GLM-5.2 leads 70.4 to 63.8.
  • The biggest single-benchmark swing is SimpleQA Verified: 34.2% for GLM-5.2 and 52% for Qwen3.6 Max Preview.
  • GLM-5.2 is cheaper at $1.40 / $4.40 per million input/output tokens, against $1.30 / $7.80 for Qwen3.6 Max Preview.
  • GLM-5.2 accepts more context: 1M tokens versus 262K.
  • GLM-5.2 has downloadable open weights; the other is API-only.

Side by side

GLM-5.2 and Qwen3.6 Max Preview specifications
GLM-5.2Qwen3.6 Max Preview
ProviderZ.ai (Zhipu)Alibaba (Qwen)
Noometry Index51.151.5
Released2026-06-132026-04-20
WeightsOpenProprietary
Context window1M262K
Max output131K66K
Input $ / M tokens$1.40$1.30
Output $ / M tokens$4.40$7.80
Results tracked5129

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5.2 leads

GLM-5.2: 51.3 (#41), Qwen3.6 Max Preview: 48.7 (#54)

Coding benchmarks
BenchmarkGLM-5.2Qwen3.6 Max Preview
SWE-bench Verified78.7%76.7%
LMArena WebDev16031482
LMArena Coding14851471
DeepSWE43.8%—
FrontierCode24.5%—
SciCode50.5%—
WeirdML70.1%—
ALE-Bench1,047—

Agentic & Tool Use Not comparable

GLM-5.2: 32.4 (#63), Qwen3.6 Max Preview: —

Agentic & Tool Use benchmarks
BenchmarkGLM-5.2Qwen3.6 Max Preview
Vending-Bench 28,3144,254
APEX-Agents45.2%—
τ²-bench Banking37.1%—
PostTrainBench31.7%—
GBAEval0%—

Reasoning Too close to call

GLM-5.2: 42.3 (#52), Qwen3.6 Max Preview: 41.7 (#53)

Reasoning benchmarks
BenchmarkGLM-5.2Qwen3.6 Max Preview
SimpleBench58.8%63%
NYT Connections (extended)74.3%74.1%
Chess Puzzles21%20%
LMArena Hard Prompts14801457
Mystery Game Puzzles19%19%
DTBench93.6%87.2%
LMCA45.8%42.5%
Epoch Capabilities Index151.78149.24
ARC-AGI-222.8%—
Kagi LLM Benchmark62.6%—
ARC-AGI-177%—
CritPt20.9%—
EBR-Bench9.5%—
Surface Evolver Bench55.6%—

Math GLM-5.2 leads

GLM-5.2: 55.7 (#43), Qwen3.6 Max Preview: 54.1 (#46)

Knowledge Too close to call

GLM-5.2: 57.1 (#40), Qwen3.6 Max Preview: 57.6 (#39)

Knowledge benchmarks
BenchmarkGLM-5.2Qwen3.6 Max Preview
GPQA Diamond91.9%87.4%
SimpleQA Verified34.2%52%
LMArena Expert14861478

Multilingual GLM-5.2 leads

GLM-5.2: 55.8 (#26), Qwen3.6 Max Preview: 54.2 (#48)

Multilingual benchmarks
BenchmarkGLM-5.2Qwen3.6 Max Preview
LMArena Non-English14591437
LMArena Chinese15191487
LMArena French14791449
LMArena Russian14661445
LMArena Spanish14771454
LMArena German1468—
LMArena Japanese1451—
LMArena Korean1445—

Instruction Following GLM-5.2 leads

GLM-5.2: 76.9 (#34), Qwen3.6 Max Preview: 75.7 (#55)

Instruction Following benchmarks
BenchmarkGLM-5.2Qwen3.6 Max Preview
LMArena Instruction Following14651438

Long Context Too close to call

GLM-5.2: 45.3 (#43), Qwen3.6 Max Preview: 44.6 (#61)

Long Context benchmarks
BenchmarkGLM-5.2Qwen3.6 Max Preview
LMArena Longer Query14791457

Writing & Preference GLM-5.2 leads

GLM-5.2: 70.4 (#21), Qwen3.6 Max Preview: 63.8 (#60)

Writing & Preference benchmarks
BenchmarkGLM-5.2Qwen3.6 Max Preview
LMArena Text14701447
LMArena Creative Writing14621435
LMArena Multi-Turn14691456
EQ-Bench Creative Writing1757—
EQ-Bench 41222—

Frequently asked questions

Is GLM-5.2 better than Qwen3.6 Max Preview?

GLM-5.2 and Qwen3.6 Max Preview score almost the same on the Noometry Index (51.1 vs 51.5), so choose on price, context window or the category you care about most.

Which is cheaper, GLM-5.2 or Qwen3.6 Max Preview?

GLM-5.2 is cheaper. It lists at $1.40 per million input tokens and $4.40 per million output tokens; Qwen3.6 Max Preview lists at $1.30 and $7.80.

Is GLM-5.2 or Qwen3.6 Max Preview better for coding?

GLM-5.2 scores higher on coding benchmarks: 51.3 versus 48.7 in the Noometry coding category.

Which has the bigger context window?

GLM-5.2 does, with 1M tokens against 262K.

How many benchmarks do GLM-5.2 and Qwen3.6 Max Preview share?

27 benchmarks have published results for both models. GLM-5.2 has 51 scored results on Noometry and Qwen3.6 Max Preview has 29.

Related comparisons

Go deeper