Model comparison

GLM-5 vs Grok 4.20 (Non-Reasoning)

Grok 4.20 (Non-Reasoning) is the stronger model overall, scoring 48.6 to 46.1 on the Noometry Index.

Last verified . 34 shared benchmarks.

GLM-5 Z.ai (Zhipu)

46.1

Rank #66 Confirmed

Grok 4.20 (Non-Reasoning) xAI

48.6

Rank #54 Confirmed

Summary

  • They share 34 benchmarks with published results for both. GLM-5 scores higher in 3 categories and Grok 4.20 (Non-Reasoning) in 6 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.20 (Non-Reasoning) leads 52.3 to 27.6.
  • The biggest single-benchmark swing is ARC-AGI-2: 4.9% for GLM-5 and 65.1% for Grok 4.20 (Non-Reasoning).
  • Both cost about the same: $1 input and $3.20 output per million tokens.
  • Grok 4.20 (Non-Reasoning) accepts more context: 1M tokens versus 205K.
  • GLM-5 has downloadable open weights; the other is API-only.

Side by side

GLM-5 and Grok 4.20 (Non-Reasoning) specifications
GLM-5Grok 4.20 (Non-Reasoning)
ProviderZ.ai (Zhipu)xAI
Noometry Index46.148.6
Released2026-02-112026-02-17
WeightsOpenProprietary
Context window205K1M
Max output131K30K
Input $ / M tokens$1$1.25
Output $ / M tokens$3.20$2.50
Results tracked4546

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5 leads

GLM-5: 49.0 (#52), Grok 4.20 (Non-Reasoning): 42.1 (#112)

Coding benchmarks
BenchmarkGLM-5Grok 4.20 (Non-Reasoning)
LMArena WebDev14341375
WeirdML48.2%52.3%
LMArena Coding14611459
ALE-Bench765.621,150
SWE-bench Verified72.1%—
SWE-bench Verified (bash only)72.8%—
SWE-bench Multilingual69.7%—

Agentic & Tool Use Grok 4.20 (Non-Reasoning) leads

GLM-5: 31.1 (#71), Grok 4.20 (Non-Reasoning): 34.4 (#46)

Agentic & Tool Use benchmarks
BenchmarkGLM-5Grok 4.20 (Non-Reasoning)
Terminal-Bench52.4%57.3%
τ²-bench Banking9.8%18%
Vending-Bench 24,4324,663
τ²-bench Airline82.5%—
τ²-bench Retail73.7%—
τ²-bench Telecom86.8%—
LMArena Search—1189

Reasoning Grok 4.20 (Non-Reasoning) leads

GLM-5: 27.6 (#116), Grok 4.20 (Non-Reasoning): 52.3 (#32)

Reasoning benchmarks
BenchmarkGLM-5Grok 4.20 (Non-Reasoning)
ARC-AGI-24.9%65.1%
Kagi LLM Benchmark75%75%
NYT Connections (extended)74.8%85.4%
ARC-AGI-144.7%89.5%
Chess Puzzles10%24%
LMArena Hard Prompts14521451
Epoch Capabilities Index145.83151.98
ForecastBench6161.4
SimpleBench53.2%—
Thematic Generalization—63.8%
DTBench—90.1%
LMCA—38.7%

Math Grok 4.20 (Non-Reasoning) leads

GLM-5: 46.4 (#71), Grok 4.20 (Non-Reasoning): 48.2 (#65)

Knowledge Too close to call

GLM-5: 52.3 (#64), Grok 4.20 (Non-Reasoning): 52.8 (#60)

Knowledge benchmarks
BenchmarkGLM-5Grok 4.20 (Non-Reasoning)
GPQA Diamond87.8%89.3%
LMArena Expert14541439
SimpleQA Verified—30.2%
Vectara Hallucination Rate10.1%—

Multimodal Not comparable

GLM-5: —, Grok 4.20 (Non-Reasoning): 33.3 (#98)

Multimodal benchmarks
BenchmarkGLM-5Grok 4.20 (Non-Reasoning)
LMArena Vision—1263
Blueprint-Bench 2—0%
LMArena Document—1416

Multilingual Too close to call

GLM-5: 53.7 (#58), Grok 4.20 (Non-Reasoning): 54.5 (#40)

Multilingual benchmarks
BenchmarkGLM-5Grok 4.20 (Non-Reasoning)
LMArena Non-English14301441
LMArena Chinese15111481
LMArena French14551476
LMArena German14451465
LMArena Japanese14161449
LMArena Korean14231417
LMArena Russian14361458
LMArena Spanish14541443

Instruction Following Too close to call

GLM-5: 75.2 (#67), Grok 4.20 (Non-Reasoning): 74.8 (#83)

Instruction Following benchmarks
BenchmarkGLM-5Grok 4.20 (Non-Reasoning)
LMArena Instruction Following14281420

Long Context Too close to call

GLM-5: 44.7 (#60), Grok 4.20 (Non-Reasoning): 45.5 (#34)

Long Context benchmarks
BenchmarkGLM-5Grok 4.20 (Non-Reasoning)
CL-bench18.7%22.2%
LMArena Longer Query14461437
CL-bench Life—11.9%

Writing & Preference Too close to call

GLM-5: 66.0 (#38), Grok 4.20 (Non-Reasoning): 65.7 (#44)

Writing & Preference benchmarks
BenchmarkGLM-5Grok 4.20 (Non-Reasoning)
LMArena Text14461451
LMArena Creative Writing14391438
EQ-Bench Creative Writing16011574
LMArena Multi-Turn14561456

Frequently asked questions

Is GLM-5 better than Grok 4.20 (Non-Reasoning)?

Grok 4.20 (Non-Reasoning) is the stronger model overall, scoring 48.6 to 46.1 on the Noometry Index.

Which is cheaper, GLM-5 or Grok 4.20 (Non-Reasoning)?

GLM-5 is cheaper. It lists at $1 per million input tokens and $3.20 per million output tokens; Grok 4.20 (Non-Reasoning) lists at $1.25 and $2.50.

Is GLM-5 or Grok 4.20 (Non-Reasoning) better for coding?

GLM-5 scores higher on coding benchmarks: 49.0 versus 42.1 in the Noometry coding category.

Which has the bigger context window?

Grok 4.20 (Non-Reasoning) does, with 1M tokens against 205K.

How many benchmarks do GLM-5 and Grok 4.20 (Non-Reasoning) share?

34 benchmarks have published results for both models. GLM-5 has 45 scored results on Noometry and Grok 4.20 (Non-Reasoning) has 46.

Related comparisons

Go deeper