Model comparison

GLM-4.6 vs Grok 4.3

Grok 4.3 is the stronger model overall, scoring 43.8 to 41.4 on the Noometry Index. GLM-4.6 costs 1.6× less per token, which makes it the better buy when Grok 4.3's lead doesn't matter for your workload.

Last verified . 21 shared benchmarks.

GLM-4.6 Z.ai (Zhipu)

41.4

Rank #135 Confirmed

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Summary

  • They share 21 benchmarks with published results for both. GLM-4.6 scores higher in 5 categories and Grok 4.3 in 4 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok 4.3 leads 52.5 to 40.2.
  • The biggest single-benchmark swing is SciCode: 38.4% for GLM-4.6 and 47.3% for Grok 4.3.
  • GLM-4.6 is cheaper at $0.60 / $2.20 per million input/output tokens, against $1.25 / $2.50 for Grok 4.3.
  • Grok 4.3 accepts more context: 1M tokens versus 205K.
  • GLM-4.6 has downloadable open weights; the other is API-only.

Side by side

GLM-4.6 and Grok 4.3 specifications
GLM-4.6Grok 4.3
ProviderZ.ai (Zhipu)xAI
Noometry Index41.443.8
Released2025-09-302026-04-17
WeightsOpenProprietary
Context window205K1M
Max output131K30K
Input $ / M tokens$0.60$1.25
Output $ / M tokens$2.20$2.50
Results tracked2940

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.3 leads

GLM-4.6: 40.1 (#148), Grok 4.3: 41.6 (#121)

Coding benchmarks
BenchmarkGLM-4.6Grok 4.3
LMArena WebDev13401357
SciCode38.4%47.3%
LMArena Coding14491415
ALE-Bench340.82944.17
SWE-bench Verified (bash only)55.4%—
WeirdML—49.9%

Agentic & Tool Use GLM-4.6 leads

GLM-4.6: 32.3 (#66), Grok 4.3: 27.7 (#99)

Agentic & Tool Use benchmarks
BenchmarkGLM-4.6Grok 4.3
Terminal-Bench24.5%—
Berkeley Function Calling Leaderboard72.4%—
GDP.pdf—8%
LMArena Search—1165
Vending-Bench 2—35.26

Reasoning Grok 4.3 leads

GLM-4.6: 23.7 (#172), Grok 4.3: 35.9 (#68)

Reasoning benchmarks
BenchmarkGLM-4.6Grok 4.3
CritPt1.1%8%
LMArena Hard Prompts14401396
Kagi LLM Benchmark47.4%—
NYT Connections (extended)—55.2%
Chess Puzzles—25%
DTBench—90.7%
LMCA—38.3%
Epoch Capabilities Index—149.16
ForecastBench—60.3

Math Grok 4.3 leads

GLM-4.6: 39.1 (#111), Grok 4.3: 46.0 (#74)

Knowledge Grok 4.3 leads

GLM-4.6: 40.2 (#124), Grok 4.3: 52.5 (#62)

Knowledge benchmarks
BenchmarkGLM-4.6Grok 4.3
LMArena Expert14311385
GPQA Diamond—88.8%
SimpleQA Verified—33.2%
Vectara Hallucination Rate9.5%—

Multimodal Not comparable

GLM-4.6: —, Grok 4.3: 31.6 (#104)

Multimodal benchmarks
BenchmarkGLM-4.6Grok 4.3
LMArena Vision—1229
Blueprint-Bench 2—0%

Multilingual GLM-4.6 leads

GLM-4.6: 53.5 (#66), Grok 4.3: 50.5 (#120)

Multilingual benchmarks
BenchmarkGLM-4.6Grok 4.3
LMArena Non-English14261385
LMArena Chinese14991422
LMArena French14591412
LMArena German14471395
LMArena Japanese13931379
LMArena Korean14001356
LMArena Russian14191399
LMArena Spanish14361398

Instruction Following GLM-4.6 leads

GLM-4.6: 74.3 (#98), Grok 4.3: 72.1 (#140)

Instruction Following benchmarks
BenchmarkGLM-4.6Grok 4.3
LMArena Instruction Following14101366

Long Context Too close to call

GLM-4.6: 43.4 (#94), Grok 4.3: 42.5 (#123)

Long Context benchmarks
BenchmarkGLM-4.6Grok 4.3
LMArena Longer Query14221393

Writing & Preference GLM-4.6 leads

GLM-4.6: 61.1 (#90), Grok 4.3: 58.5 (#118)

Writing & Preference benchmarks
BenchmarkGLM-4.6Grok 4.3
LMArena Text14401397
LMArena Creative Writing14111380
LMArena Multi-Turn14271406
EQ-Bench Creative Writing1411—
EQ-Bench 4—1075

Frequently asked questions

Is GLM-4.6 better than Grok 4.3?

Grok 4.3 is the stronger model overall, scoring 43.8 to 41.4 on the Noometry Index. GLM-4.6 costs 1.6× less per token, which makes it the better buy when Grok 4.3's lead doesn't matter for your workload.

Which is cheaper, GLM-4.6 or Grok 4.3?

GLM-4.6 is cheaper. It lists at $0.60 per million input tokens and $2.20 per million output tokens; Grok 4.3 lists at $1.25 and $2.50.

Is GLM-4.6 or Grok 4.3 better for coding?

Grok 4.3 scores higher on coding benchmarks: 41.6 versus 40.1 in the Noometry coding category.

Which has the bigger context window?

Grok 4.3 does, with 1M tokens against 205K.

How many benchmarks do GLM-4.6 and Grok 4.3 share?

21 benchmarks have published results for both models. GLM-4.6 has 29 scored results on Noometry and Grok 4.3 has 40.

Related comparisons

Go deeper