Model comparison

GLM-5.2 vs Inkling

GLM-5.2 is the stronger model overall, scoring 51.1 to 44.1 on the Noometry Index.

Last verified . 40 shared benchmarks.

GLM-5.2 Z.ai (Zhipu)

51.1

Rank #44 Confirmed

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Summary

  • They share 40 benchmarks with published results for both. GLM-5.2 scores higher in 9 categories and Inkling in 0 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where GLM-5.2 leads 55.7 to 31.3.
  • The biggest single-benchmark swing is WeirdML: 70.1% for GLM-5.2 and 32.3% for Inkling.
  • GLM-5.2 is cheaper at $1.40 / $4.40 per million input/output tokens, against $1.87 / $4.68 for Inkling.
  • GLM-5.2 accepts more context: 1M tokens versus 66K.

Side by side

GLM-5.2 and Inkling specifications
GLM-5.2Inkling
ProviderZ.ai (Zhipu)Thinking Machines Lab
Noometry Index51.144.1
Released2026-06-132026-07-15
WeightsOpenOpen
Context window1M66K
Max output131K66K
Input $ / M tokens$1.40$1.87
Output $ / M tokens$4.40$4.68
Results tracked5141

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5.2 leads

GLM-5.2: 51.3 (#41), Inkling: 34.5 (#234)

Coding benchmarks
BenchmarkGLM-5.2Inkling
FrontierCode24.5%14%
LMArena WebDev16031413
SciCode50.5%47%
WeirdML70.1%32.3%
LMArena Coding14851464
ALE-Bench1,047946
SWE-bench Verified78.7%—
DeepSWE43.8%—
FrontierSWE—4.1%

Agentic & Tool Use GLM-5.2 leads

GLM-5.2: 32.4 (#63), Inkling: 29.6 (#85)

Agentic & Tool Use benchmarks
BenchmarkGLM-5.2Inkling
APEX-Agents45.2%33.8%
τ²-bench Banking37.1%25%
PostTrainBench31.7%—
GBAEval0%—
Vending-Bench 28,314—

Reasoning GLM-5.2 leads

GLM-5.2: 42.3 (#52), Inkling: 40.4 (#56)

Reasoning benchmarks
BenchmarkGLM-5.2Inkling
ARC-AGI-222.8%36.5%
SimpleBench58.8%50%
ARC-AGI-177%79.5%
CritPt20.9%5.4%
Chess Puzzles21%21%
LMArena Hard Prompts14801451
DTBench93.6%87.5%
LMCA45.8%37.6%
Epoch Capabilities Index151.78148.54
Kagi LLM Benchmark62.6%—
NYT Connections (extended)74.3%—
EBR-Bench9.5%—
Mystery Game Puzzles19%—
Surface Evolver Bench55.6%—

Math GLM-5.2 leads

GLM-5.2: 55.7 (#43), Inkling: 31.3 (#225)

Math benchmarks
BenchmarkGLM-5.2Inkling
FrontierMath (Tiers 1-3)59.2%33.3%
FrontierMath Tier 429.3%4.9%
OTIS Mock AIME 2024-202586.4%88.9%
ProofBench35%0%
LMArena Math14821479
MathArena Final-Answer Competitions67.6%—

Knowledge GLM-5.2 leads

GLM-5.2: 57.1 (#40), Inkling: 55.1 (#49)

Knowledge benchmarks
BenchmarkGLM-5.2Inkling
GPQA Diamond91.9%88.3%
SimpleQA Verified34.2%40.3%
LMArena Expert14861465

Multilingual GLM-5.2 leads

GLM-5.2: 55.8 (#26), Inkling: 54.0 (#52)

Multilingual benchmarks
BenchmarkGLM-5.2Inkling
LMArena Non-English14591434
LMArena Chinese15191490
LMArena French14791458
LMArena German14681446
LMArena Japanese14511429
LMArena Korean14451404
LMArena Russian14661429
LMArena Spanish14771448

Instruction Following GLM-5.2 leads

GLM-5.2: 76.9 (#34), Inkling: 75.1 (#71)

Instruction Following benchmarks
BenchmarkGLM-5.2Inkling
LMArena Instruction Following14651426

Long Context GLM-5.2 leads

GLM-5.2: 45.3 (#43), Inkling: 43.8 (#86)

Long Context benchmarks
BenchmarkGLM-5.2Inkling
LMArena Longer Query14791434

Writing & Preference GLM-5.2 leads

GLM-5.2: 70.4 (#21), Inkling: 65.2 (#51)

Writing & Preference benchmarks
BenchmarkGLM-5.2Inkling
LMArena Text14701441
LMArena Creative Writing14621387
EQ-Bench Creative Writing17571611
EQ-Bench 412221226
LMArena Multi-Turn14691436

Frequently asked questions

Is GLM-5.2 better than Inkling?

GLM-5.2 is the stronger model overall, scoring 51.1 to 44.1 on the Noometry Index.

Which is cheaper, GLM-5.2 or Inkling?

GLM-5.2 is cheaper. It lists at $1.40 per million input tokens and $4.40 per million output tokens; Inkling lists at $1.87 and $4.68.

Is GLM-5.2 or Inkling better for coding?

GLM-5.2 scores higher on coding benchmarks: 51.3 versus 34.5 in the Noometry coding category.

Which has the bigger context window?

GLM-5.2 does, with 1M tokens against 66K.

How many benchmarks do GLM-5.2 and Inkling share?

40 benchmarks have published results for both models. GLM-5.2 has 51 scored results on Noometry and Inkling has 41.

Related comparisons

Go deeper