Model comparison

GLM-5.2 vs GLM-5.3-Flash

GLM-5.2 and GLM-5.3-Flash score almost the same on the Noometry Index (51.1 vs 51.8), so choose on price, context window or the category you care about most.

Last verified . 35 shared benchmarks.

GLM-5.2 Z.ai (Zhipu)

51.1

Rank #44 Confirmed

GLM-5.3-Flash Z.ai (Zhipu)

51.8

Rank #41 Confirmed

Summary

  • They share 35 benchmarks with published results for both. GLM-5.2 scores higher in 2 categories and GLM-5.3-Flash in 7 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where GLM-5.3-Flash leads 48.0 to 42.3.
  • The biggest single-benchmark swing is ARC-AGI-2: 22.8% for GLM-5.2 and 65.8% for GLM-5.3-Flash.
  • GLM-5.3-Flash is cheaper at $0.15 / $0.50 per million input/output tokens, against $1.40 / $4.40 for GLM-5.2.

Side by side

GLM-5.2 and GLM-5.3-Flash specifications
GLM-5.2GLM-5.3-Flash
ProviderZ.ai (Zhipu)Z.ai (Zhipu)
Noometry Index51.151.8
Released2026-06-132026-08-20
WeightsOpenOpen
Context window1M1M
Max output131K131K
Input $ / M tokens$1.40$0.15
Output $ / M tokens$4.40$0.50
Results tracked5140

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5.3-Flash leads

GLM-5.2: 51.3 (#41), GLM-5.3-Flash: 53.1 (#31)

Coding benchmarks
BenchmarkGLM-5.2GLM-5.3-Flash
DeepSWE43.8%63.4%
FrontierCode24.5%31.8%
LMArena WebDev16031609
SciCode50.5%51.6%
LMArena Coding14851508
ALE-Bench1,047303.55
SWE-bench Verified78.7%—
CursorBench—36.8%
FrontierSWE—18.1%
WeirdML70.1%—

Agentic & Tool Use GLM-5.3-Flash leads

GLM-5.2: 32.4 (#63), GLM-5.3-Flash: 34.2 (#47)

Agentic & Tool Use benchmarks
BenchmarkGLM-5.2GLM-5.3-Flash
APEX-Agents45.2%52.8%
τ²-bench Banking37.1%—
PostTrainBench31.7%—
GBAEval0%—
GDP.pdf—14%
Vending-Bench 28,314—

Reasoning GLM-5.3-Flash leads

GLM-5.2: 42.3 (#52), GLM-5.3-Flash: 48.0 (#42)

Reasoning benchmarks
BenchmarkGLM-5.2GLM-5.3-Flash
ARC-AGI-222.8%65.8%
ARC-AGI-177%91%
CritPt20.9%15.4%
Chess Puzzles21%14%
LMArena Hard Prompts14801491
Mystery Game Puzzles19%8%
Surface Evolver Bench55.6%52.5%
Epoch Capabilities Index151.78151.88
SimpleBench58.8%—
Kagi LLM Benchmark62.6%—
NYT Connections (extended)74.3%—
EBR-Bench9.5%—
DTBench93.6%—
LMCA45.8%—
Bench to the Future 3—0.15

Math GLM-5.2 leads

GLM-5.2: 55.7 (#43), GLM-5.3-Flash: 53.3 (#47)

Math benchmarks
BenchmarkGLM-5.2GLM-5.3-Flash
FrontierMath (Tiers 1-3)59.2%55.8%
FrontierMath Tier 429.3%17.1%
OTIS Mock AIME 2024-202586.4%93.9%
ProofBench35%21%
LMArena Math14821500
MathArena Final-Answer Competitions67.6%—

Knowledge GLM-5.3-Flash leads

GLM-5.2: 57.1 (#40), GLM-5.3-Flash: 58.4 (#36)

Knowledge benchmarks
BenchmarkGLM-5.2GLM-5.3-Flash
GPQA Diamond91.9%90.2%
LMArena Expert14861513
SimpleQA Verified34.2%—

Multimodal Not comparable

GLM-5.2: —, GLM-5.3-Flash: 42.8 (#27)

Multimodal benchmarks
BenchmarkGLM-5.2GLM-5.3-Flash
LMArena Vision—1296

Multilingual Too close to call

GLM-5.2: 55.8 (#26), GLM-5.3-Flash: 56.0 (#25)

Multilingual benchmarks
BenchmarkGLM-5.2GLM-5.3-Flash
LMArena Non-English14591462
LMArena Chinese15191527
LMArena French14791496
LMArena German14681470
LMArena Japanese14511429
LMArena Korean14451446
LMArena Russian14661469
LMArena Spanish14771471

Instruction Following Too close to call

GLM-5.2: 76.9 (#34), GLM-5.3-Flash: 77.5 (#20)

Instruction Following benchmarks
BenchmarkGLM-5.2GLM-5.3-Flash
LMArena Instruction Following14651478

Long Context Too close to call

GLM-5.2: 45.3 (#43), GLM-5.3-Flash: 45.4 (#39)

Long Context benchmarks
BenchmarkGLM-5.2GLM-5.3-Flash
LMArena Longer Query14791482

Writing & Preference GLM-5.2 leads

GLM-5.2: 70.4 (#21), GLM-5.3-Flash: 65.3 (#50)

Writing & Preference benchmarks
BenchmarkGLM-5.2GLM-5.3-Flash
LMArena Text14701471
LMArena Creative Writing14621442
LMArena Multi-Turn14691467
EQ-Bench Creative Writing1757—
EQ-Bench 41222—

Frequently asked questions

Is GLM-5.2 better than GLM-5.3-Flash?

GLM-5.2 and GLM-5.3-Flash score almost the same on the Noometry Index (51.1 vs 51.8), so choose on price, context window or the category you care about most.

Which is cheaper, GLM-5.2 or GLM-5.3-Flash?

GLM-5.3-Flash is cheaper. It lists at $0.15 per million input tokens and $0.50 per million output tokens; GLM-5.2 lists at $1.40 and $4.40.

Is GLM-5.2 or GLM-5.3-Flash better for coding?

GLM-5.3-Flash scores higher on coding benchmarks: 53.1 versus 51.3 in the Noometry coding category.

Which has the bigger context window?

Both accept 1M tokens.

How many benchmarks do GLM-5.2 and GLM-5.3-Flash share?

35 benchmarks have published results for both models. GLM-5.2 has 51 scored results on Noometry and GLM-5.3-Flash has 40.

Related comparisons

Go deeper