Model comparison

GLM-5.1 vs GLM-5.2

GLM-5.2 is the stronger model overall, scoring 51.1 to 47.8 on the Noometry Index.

Last verified . 37 shared benchmarks.

GLM-5.1 Z.ai (Zhipu)

47.8

Rank #59 Confirmed

GLM-5.2 Z.ai (Zhipu)

51.1

Rank #44 Confirmed

Summary

  • They share 37 benchmarks with published results for both. GLM-5.1 scores higher in 0 categories and GLM-5.2 in 9 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where GLM-5.2 leads 32.4 to 24.9.
  • The biggest single-benchmark swing is FrontierMath (Tiers 1-3): 36.8% for GLM-5.1 and 59.2% for GLM-5.2.
  • Both cost about the same: $1.40 input and $4.40 output per million tokens.
  • GLM-5.2 accepts more context: 1M tokens versus 200K.

Side by side

GLM-5.1 and GLM-5.2 specifications
GLM-5.1GLM-5.2
ProviderZ.ai (Zhipu)Z.ai (Zhipu)
Noometry Index47.851.1
Released2026-04-072026-06-13
WeightsOpenOpen
Context window200K1M
Max output131K131K
Input $ / M tokens$1.40$1.40
Output $ / M tokens$4.40$4.40
Results tracked4151

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5.2 leads

GLM-5.1: 48.7 (#55), GLM-5.2: 51.3 (#41)

Coding benchmarks
BenchmarkGLM-5.1GLM-5.2
SWE-bench Verified74.2%78.7%
LMArena WebDev15081603
SciCode43.8%50.5%
WeirdML57.1%70.1%
LMArena Coding14851485
ALE-Bench887.11,047
DeepSWE—43.8%
FrontierCode—24.5%

Agentic & Tool Use GLM-5.2 leads

GLM-5.1: 24.9 (#113), GLM-5.2: 32.4 (#63)

Agentic & Tool Use benchmarks
BenchmarkGLM-5.1GLM-5.2
APEX-Agents40.9%45.2%
GBAEval0%0%
Vending-Bench 25,6348,314
τ²-bench Banking—37.1%
PostTrainBench—31.7%
ExploitBench18.1%—

Reasoning GLM-5.2 leads

GLM-5.1: 39.1 (#60), GLM-5.2: 42.3 (#52)

Reasoning benchmarks
BenchmarkGLM-5.1GLM-5.2
SimpleBench55.1%58.8%
NYT Connections (extended)77.7%74.3%
CritPt4.6%20.9%
Chess Puzzles19%21%
LMArena Hard Prompts14721480
Epoch Capabilities Index149.84151.78
ARC-AGI-2—22.8%
Kagi LLM Benchmark—62.6%
ARC-AGI-1—77%
Thematic Generalization69.8%—
EBR-Bench—9.5%
Mystery Game Puzzles—19%
DTBench—93.6%
LMCA—45.8%
Surface Evolver Bench—55.6%

Math GLM-5.2 leads

GLM-5.1: 49.7 (#60), GLM-5.2: 55.7 (#43)

Knowledge GLM-5.2 leads

GLM-5.1: 54.9 (#50), GLM-5.2: 57.1 (#40)

Knowledge benchmarks
BenchmarkGLM-5.1GLM-5.2
GPQA Diamond89.9%91.9%
SimpleQA Verified34%34.2%
LMArena Expert14761486

Multilingual Too close to call

GLM-5.1: 55.0 (#36), GLM-5.2: 55.8 (#26)

Multilingual benchmarks
BenchmarkGLM-5.1GLM-5.2
LMArena Non-English14471459
LMArena Chinese15151519
LMArena French14741479
LMArena German14651468
LMArena Japanese14341451
LMArena Korean14181445
LMArena Russian14541466
LMArena Spanish14691477

Instruction Following Too close to call

GLM-5.1: 76.3 (#42), GLM-5.2: 76.9 (#34)

Instruction Following benchmarks
BenchmarkGLM-5.1GLM-5.2
LMArena Instruction Following14511465

Long Context Too close to call

GLM-5.1: 44.9 (#53), GLM-5.2: 45.3 (#43)

Long Context benchmarks
BenchmarkGLM-5.1GLM-5.2
LMArena Longer Query14661479

Writing & Preference GLM-5.2 leads

GLM-5.1: 66.9 (#31), GLM-5.2: 70.4 (#21)

Writing & Preference benchmarks
BenchmarkGLM-5.1GLM-5.2
LMArena Text14611470
LMArena Creative Writing14531462
EQ-Bench Creative Writing15921757
LMArena Multi-Turn14721469
EQ-Bench 4—1222

Frequently asked questions

Is GLM-5.1 better than GLM-5.2?

GLM-5.2 is the stronger model overall, scoring 51.1 to 47.8 on the Noometry Index.

Which is cheaper, GLM-5.1 or GLM-5.2?

GLM-5.2 is cheaper. It lists at $1.40 per million input tokens and $4.40 per million output tokens; GLM-5.1 lists at $1.40 and $4.40.

Is GLM-5.1 or GLM-5.2 better for coding?

GLM-5.2 scores higher on coding benchmarks: 51.3 versus 48.7 in the Noometry coding category.

Which has the bigger context window?

GLM-5.2 does, with 1M tokens against 200K.

How many benchmarks do GLM-5.1 and GLM-5.2 share?

37 benchmarks have published results for both models. GLM-5.1 has 41 scored results on Noometry and GLM-5.2 has 51.

Related comparisons

Go deeper