Model comparison

GLM-5.2 vs o3-mini

GLM-5.2 is the stronger model overall, scoring 51.1 to 36.7 on the Noometry Index.

Last verified . 33 shared benchmarks.

GLM-5.2 Z.ai (Zhipu)

51.1

Rank #44 Confirmed

o3-mini OpenAI

36.7

Rank #212 Confirmed

Summary

  • They share 33 benchmarks with published results for both. GLM-5.2 scores higher in 9 categories and o3-mini in 0 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where GLM-5.2 leads 55.7 to 28.1.
  • The biggest single-benchmark swing is ARC-AGI-1: 77% for GLM-5.2 and 34.5% for o3-mini.
  • o3-mini is cheaper at $1.10 / $4.40 per million input/output tokens, against $1.40 / $4.40 for GLM-5.2.
  • GLM-5.2 accepts more context: 1M tokens versus 200K.
  • GLM-5.2 has downloadable open weights; the other is API-only.

Side by side

GLM-5.2 and o3-mini specifications
GLM-5.2o3-mini
ProviderZ.ai (Zhipu)OpenAI
Noometry Index51.136.7
Released2026-06-132024-12-20
WeightsOpenProprietary
Context window1M200K
Max output131K100K
Input $ / M tokens$1.40$1.10
Output $ / M tokens$4.40$4.40
Results tracked5151

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5.2 leads

GLM-5.2: 51.3 (#41), o3-mini: 40.8 (#132)

Coding benchmarks
BenchmarkGLM-5.2o3-mini
SciCode50.5%39.8%
WeirdML70.1%43.7%
LMArena Coding14851378
SWE-bench Verified78.7%—
DeepSWE43.8%—
FrontierCode24.5%—
Aider Polyglot—60.4%
LMArena WebDev1603—
GSO—1.3%
LiveBench Coding—82.7%
CadEval—54%
ALE-Bench1,047—

Agentic & Tool Use GLM-5.2 leads

GLM-5.2: 32.4 (#63), o3-mini: 29.6 (#84)

Agentic & Tool Use benchmarks
BenchmarkGLM-5.2o3-mini
APEX-Agents45.2%—
τ²-bench Banking37.1%—
Cybench—22.5%
PostTrainBench31.7%—
GBAEval0%—
Vending-Bench 28,314—

Reasoning GLM-5.2 leads

GLM-5.2: 42.3 (#52), o3-mini: 16.3 (#305)

Reasoning benchmarks
BenchmarkGLM-5.2o3-mini
ARC-AGI-222.8%3%
SimpleBench58.8%22.8%
ARC-AGI-177%34.5%
CritPt20.9%0.3%
Chess Puzzles21%17%
LMArena Hard Prompts14801366
Mystery Game Puzzles19%7%
DTBench93.6%68.8%
LMCA45.8%19%
Epoch Capabilities Index151.78140.34
Kagi LLM Benchmark62.6%—
NYT Connections (extended)74.3%—
EBR-Bench9.5%—
LiveBench Reasoning—89.6%
LiveBench Data Analysis—70.6%
Surface Evolver Bench55.6%—
ForecastBench—59.6
LiveBench—75.9%

Math GLM-5.2 leads

GLM-5.2: 55.7 (#43), o3-mini: 28.1 (#244)

Knowledge GLM-5.2 leads

GLM-5.2: 57.1 (#40), o3-mini: 38.3 (#146)

Knowledge benchmarks
BenchmarkGLM-5.2o3-mini
GPQA Diamond91.9%77%
SimpleQA Verified34.2%15.3%
LMArena Expert14861364
Confabulations—17.9%

Multilingual GLM-5.2 leads

GLM-5.2: 55.8 (#26), o3-mini: 45.7 (#164)

Multilingual benchmarks
BenchmarkGLM-5.2o3-mini
LMArena Non-English14591319
LMArena Chinese15191379
LMArena French14791334
LMArena German14681303
LMArena Japanese14511286
LMArena Korean14451314
LMArena Russian14661304
LMArena Spanish14771321

Instruction Following GLM-5.2 leads

GLM-5.2: 76.9 (#34), o3-mini: 75.1 (#72)

Instruction Following benchmarks
BenchmarkGLM-5.2o3-mini
LMArena Instruction Following14651337
LiveBench Instruction Following—84.4%

Long Context GLM-5.2 leads

GLM-5.2: 45.3 (#43), o3-mini: 33.8 (#256)

Long Context benchmarks
BenchmarkGLM-5.2o3-mini
LMArena Longer Query14791343
Fiction.LiveBench—50%

Writing & Preference GLM-5.2 leads

GLM-5.2: 70.4 (#21), o3-mini: 50.3 (#182)

Writing & Preference benchmarks
BenchmarkGLM-5.2o3-mini
LMArena Text14701337
LMArena Creative Writing14621286
LMArena Multi-Turn14691320
Short-Story Creative Writing—61.7%
EQ-Bench Creative Writing1757—
EQ-Bench 41222—
LiveBench Language—50.7%

Frequently asked questions

Is GLM-5.2 better than o3-mini?

GLM-5.2 is the stronger model overall, scoring 51.1 to 36.7 on the Noometry Index.

Which is cheaper, GLM-5.2 or o3-mini?

o3-mini is cheaper. It lists at $1.10 per million input tokens and $4.40 per million output tokens; GLM-5.2 lists at $1.40 and $4.40.

Is GLM-5.2 or o3-mini better for coding?

GLM-5.2 scores higher on coding benchmarks: 51.3 versus 40.8 in the Noometry coding category.

Which has the bigger context window?

GLM-5.2 does, with 1M tokens against 200K.

How many benchmarks do GLM-5.2 and o3-mini share?

33 benchmarks have published results for both models. GLM-5.2 has 51 scored results on Noometry and o3-mini has 51.

Related comparisons

Go deeper