Model comparison

GLM-5.2 vs GLM-5.3

GLM-5.3 is the stronger model overall, scoring 54.8 to 51.1 on the Noometry Index.

Last verified . 39 shared benchmarks.

GLM-5.2 Z.ai (Zhipu)

51.1

Rank #44 Confirmed

GLM-5.3 Z.ai (Zhipu)

54.8

Rank #26 Confirmed

Summary

  • They share 39 benchmarks with published results for both. GLM-5.2 scores higher in 1 category and GLM-5.3 in 8 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in coding, where GLM-5.3 leads 59.5 to 51.3.
  • The biggest single-benchmark swing is DeepSWE: 43.8% for GLM-5.2 and 69% for GLM-5.3.
  • Both cost about the same: $1.40 input and $4.40 output per million tokens.

Side by side

GLM-5.2 and GLM-5.3 specifications
GLM-5.2GLM-5.3
ProviderZ.ai (Zhipu)Z.ai (Zhipu)
Noometry Index51.154.8
Released2026-06-132026-08-14
WeightsOpenOpen
Context window1M1M
Max output131K131K
Input $ / M tokens$1.40$1.40
Output $ / M tokens$4.40$4.40
Results tracked5142

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5.3 leads

GLM-5.2: 51.3 (#41), GLM-5.3: 59.5 (#14)

Coding benchmarks
BenchmarkGLM-5.2GLM-5.3
DeepSWE43.8%69%
FrontierCode24.5%40.1%
LMArena WebDev16031622
SciCode50.5%59%
WeirdML70.1%75.4%
LMArena Coding14851496
ALE-Bench1,0471,317
SWE-bench Verified78.7%—
CursorBench—42.6%
FrontierSWE—30.2%

Agentic & Tool Use GLM-5.3 leads

GLM-5.2: 32.4 (#63), GLM-5.3: 36.4 (#38)

Agentic & Tool Use benchmarks
BenchmarkGLM-5.2GLM-5.3
APEX-Agents45.2%56.6%
Vending-Bench 28,3148,164
τ²-bench Banking37.1%—
PostTrainBench31.7%—
GBAEval0%—

Reasoning GLM-5.3 leads

GLM-5.2: 42.3 (#52), GLM-5.3: 46.1 (#46)

Reasoning benchmarks
BenchmarkGLM-5.2GLM-5.3
NYT Connections (extended)74.3%74.2%
CritPt20.9%19.1%
Chess Puzzles21%21%
LMArena Hard Prompts14801489
Mystery Game Puzzles19%33%
DTBench93.6%87.7%
LMCA45.8%55.5%
Epoch Capabilities Index151.78155.61
ARC-AGI-222.8%—
SimpleBench58.8%—
Kagi LLM Benchmark62.6%—
ARC-AGI-177%—
EBR-Bench9.5%—
Surface Evolver Bench55.6%—
Bench to the Future 3—0.15

Math GLM-5.3 leads

GLM-5.2: 55.7 (#43), GLM-5.3: 62.3 (#33)

Math benchmarks
BenchmarkGLM-5.2GLM-5.3
FrontierMath (Tiers 1-3)59.2%68.8%
FrontierMath Tier 429.3%29.3%
OTIS Mock AIME 2024-202586.4%91.1%
ProofBench35%49%
LMArena Math14821489
MathArena Final-Answer Competitions67.6%—

Knowledge GLM-5.3 leads

GLM-5.2: 57.1 (#40), GLM-5.3: 58.3 (#37)

Knowledge benchmarks
BenchmarkGLM-5.2GLM-5.3
GPQA Diamond91.9%90.9%
SimpleQA Verified34.2%41%
LMArena Expert14861516

Multilingual Too close to call

GLM-5.2: 55.8 (#26), GLM-5.3: 55.7 (#28)

Multilingual benchmarks
BenchmarkGLM-5.2GLM-5.3
LMArena Non-English14591457
LMArena Chinese15191528
LMArena French14791499
LMArena German14681499
LMArena Japanese14511453
LMArena Korean14451472
LMArena Russian14661463
LMArena Spanish14771460

Instruction Following Too close to call

GLM-5.2: 76.9 (#34), GLM-5.3: 77.5 (#23)

Instruction Following benchmarks
BenchmarkGLM-5.2GLM-5.3
LMArena Instruction Following14651477

Long Context Too close to call

GLM-5.2: 45.3 (#43), GLM-5.3: 45.4 (#41)

Long Context benchmarks
BenchmarkGLM-5.2GLM-5.3
LMArena Longer Query14791482

Writing & Preference GLM-5.3 leads

GLM-5.2: 70.4 (#21), GLM-5.3: 75.7 (#6)

Writing & Preference benchmarks
BenchmarkGLM-5.2GLM-5.3
LMArena Text14701471
LMArena Creative Writing14621457
EQ-Bench Creative Writing17572075
LMArena Multi-Turn14691472
EQ-Bench 41222—

Frequently asked questions

Is GLM-5.2 better than GLM-5.3?

GLM-5.3 is the stronger model overall, scoring 54.8 to 51.1 on the Noometry Index.

Which is cheaper, GLM-5.2 or GLM-5.3?

GLM-5.3 is cheaper. It lists at $1.40 per million input tokens and $4.40 per million output tokens; GLM-5.2 lists at $1.40 and $4.40.

Is GLM-5.2 or GLM-5.3 better for coding?

GLM-5.3 scores higher on coding benchmarks: 59.5 versus 51.3 in the Noometry coding category.

Which has the bigger context window?

Both accept 1M tokens.

How many benchmarks do GLM-5.2 and GLM-5.3 share?

39 benchmarks have published results for both models. GLM-5.2 has 51 scored results on Noometry and GLM-5.3 has 42.

Related comparisons

Go deeper