Model comparison

GLM-5.1 vs Grok 4.20 (Non-Reasoning)

GLM-5.1 and Grok 4.20 (Non-Reasoning) score almost the same on the Noometry Index (47.8 vs 48.6), so choose on price, context window or the category you care about most.

Last verified . 31 shared benchmarks.

GLM-5.1 Z.ai (Zhipu)

47.8

Rank #59 Confirmed

Grok 4.20 (Non-Reasoning) xAI

48.6

Rank #54 Confirmed

Summary

  • They share 31 benchmarks with published results for both. GLM-5.1 scores higher in 6 categories and Grok 4.20 (Non-Reasoning) in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.20 (Non-Reasoning) leads 52.3 to 39.1.
  • The biggest single-benchmark swing is ProofBench: 22.2% for GLM-5.1 and 14% for Grok 4.20 (Non-Reasoning).
  • Grok 4.20 (Non-Reasoning) is cheaper at $1.25 / $2.50 per million input/output tokens, against $1.40 / $4.40 for GLM-5.1.
  • Grok 4.20 (Non-Reasoning) accepts more context: 1M tokens versus 200K.
  • GLM-5.1 has downloadable open weights; the other is API-only.

Side by side

GLM-5.1 and Grok 4.20 (Non-Reasoning) specifications
GLM-5.1Grok 4.20 (Non-Reasoning)
ProviderZ.ai (Zhipu)xAI
Noometry Index47.848.6
Released2026-04-072026-02-17
WeightsOpenProprietary
Context window200K1M
Max output131K30K
Input $ / M tokens$1.40$1.25
Output $ / M tokens$4.40$2.50
Results tracked4146

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5.1 leads

GLM-5.1: 48.7 (#55), Grok 4.20 (Non-Reasoning): 42.1 (#112)

Coding benchmarks
BenchmarkGLM-5.1Grok 4.20 (Non-Reasoning)
LMArena WebDev15081375
WeirdML57.1%52.3%
LMArena Coding14851459
ALE-Bench887.11,150
SWE-bench Verified74.2%—
SciCode43.8%—

Agentic & Tool Use Grok 4.20 (Non-Reasoning) leads

GLM-5.1: 24.9 (#113), Grok 4.20 (Non-Reasoning): 34.4 (#46)

Agentic & Tool Use benchmarks
BenchmarkGLM-5.1Grok 4.20 (Non-Reasoning)
Vending-Bench 25,6344,663
Terminal-Bench—57.3%
APEX-Agents40.9%—
τ²-bench Banking—18%
ExploitBench18.1%—
GBAEval0%—
LMArena Search—1189

Reasoning Grok 4.20 (Non-Reasoning) leads

GLM-5.1: 39.1 (#60), Grok 4.20 (Non-Reasoning): 52.3 (#32)

Reasoning benchmarks
BenchmarkGLM-5.1Grok 4.20 (Non-Reasoning)
NYT Connections (extended)77.7%85.4%
Chess Puzzles19%24%
Thematic Generalization69.8%63.8%
LMArena Hard Prompts14721451
Epoch Capabilities Index149.84151.98
ARC-AGI-2—65.1%
SimpleBench55.1%—
Kagi LLM Benchmark—75%
ARC-AGI-1—89.5%
CritPt4.6%—
DTBench—90.1%
LMCA—38.7%
ForecastBench—61.4

Math GLM-5.1 leads

GLM-5.1: 49.7 (#60), Grok 4.20 (Non-Reasoning): 48.2 (#65)

Math benchmarks
BenchmarkGLM-5.1Grok 4.20 (Non-Reasoning)
FrontierMath (Tiers 1-3)36.8%44.9%
OTIS Mock AIME 2024-202593.3%92.2%
ProofBench22.2%14%
LMArena Math14731455
FrontierMath Tier 4—17.1%
MathArena Final-Answer Competitions67.1%—
FrontierMath (Feb 2025 set)33.4%—
FrontierMath Tier 4 (v1)12.5%—

Knowledge GLM-5.1 leads

GLM-5.1: 54.9 (#50), Grok 4.20 (Non-Reasoning): 52.8 (#60)

Knowledge benchmarks
BenchmarkGLM-5.1Grok 4.20 (Non-Reasoning)
GPQA Diamond89.9%89.3%
SimpleQA Verified34%30.2%
LMArena Expert14761439

Multimodal Not comparable

GLM-5.1: —, Grok 4.20 (Non-Reasoning): 33.3 (#98)

Multimodal benchmarks
BenchmarkGLM-5.1Grok 4.20 (Non-Reasoning)
LMArena Vision—1263
Blueprint-Bench 2—0%
LMArena Document—1416

Multilingual Too close to call

GLM-5.1: 55.0 (#36), Grok 4.20 (Non-Reasoning): 54.5 (#40)

Multilingual benchmarks
BenchmarkGLM-5.1Grok 4.20 (Non-Reasoning)
LMArena Non-English14471441
LMArena Chinese15151481
LMArena French14741476
LMArena German14651465
LMArena Japanese14341449
LMArena Korean14181417
LMArena Russian14541458
LMArena Spanish14691443

Instruction Following GLM-5.1 leads

GLM-5.1: 76.3 (#42), Grok 4.20 (Non-Reasoning): 74.8 (#83)

Instruction Following benchmarks
BenchmarkGLM-5.1Grok 4.20 (Non-Reasoning)
LMArena Instruction Following14511420

Long Context Too close to call

GLM-5.1: 44.9 (#53), Grok 4.20 (Non-Reasoning): 45.5 (#34)

Long Context benchmarks
BenchmarkGLM-5.1Grok 4.20 (Non-Reasoning)
LMArena Longer Query14661437
CL-bench—22.2%
CL-bench Life—11.9%

Writing & Preference GLM-5.1 leads

GLM-5.1: 66.9 (#31), Grok 4.20 (Non-Reasoning): 65.7 (#44)

Writing & Preference benchmarks
BenchmarkGLM-5.1Grok 4.20 (Non-Reasoning)
LMArena Text14611451
LMArena Creative Writing14531438
EQ-Bench Creative Writing15921574
LMArena Multi-Turn14721456

Frequently asked questions

Is GLM-5.1 better than Grok 4.20 (Non-Reasoning)?

GLM-5.1 and Grok 4.20 (Non-Reasoning) score almost the same on the Noometry Index (47.8 vs 48.6), so choose on price, context window or the category you care about most.

Which is cheaper, GLM-5.1 or Grok 4.20 (Non-Reasoning)?

Grok 4.20 (Non-Reasoning) is cheaper. It lists at $1.25 per million input tokens and $2.50 per million output tokens; GLM-5.1 lists at $1.40 and $4.40.

Is GLM-5.1 or Grok 4.20 (Non-Reasoning) better for coding?

GLM-5.1 scores higher on coding benchmarks: 48.7 versus 42.1 in the Noometry coding category.

Which has the bigger context window?

Grok 4.20 (Non-Reasoning) does, with 1M tokens against 200K.

How many benchmarks do GLM-5.1 and Grok 4.20 (Non-Reasoning) share?

31 benchmarks have published results for both models. GLM-5.1 has 41 scored results on Noometry and Grok 4.20 (Non-Reasoning) has 46.

Related comparisons

Go deeper