Model comparison

GLM-4.7-Flash vs Grok-3 mini

Grok-3 mini is the stronger model overall, scoring 41.2 to 38.8 on the Noometry Index.

Last verified . 18 shared benchmarks.

GLM-4.7-Flash Z.ai (Zhipu)

38.8

Rank #180 Confirmed

Grok-3 mini xAI

41.2

Rank #141 Confirmed

Summary

  • They share 18 benchmarks with published results for both. GLM-4.7-Flash scores higher in 1 category and Grok-3 mini in 7 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok-3 mini leads 46.4 to 35.5.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 58.3% for GLM-4.7-Flash and 77.8% for Grok-3 mini.
  • GLM-4.7-Flash has downloadable open weights; the other is API-only.

Side by side

GLM-4.7-Flash and Grok-3 mini specifications
GLM-4.7-FlashGrok-3 mini
ProviderZ.ai (Zhipu)xAI
Noometry Index38.841.2
Released2026-01-192025-04-09
WeightsOpenProprietary
Context window200K—
Max output131K—
Input $ / M tokens$0.06—
Output $ / M tokens$0.40—
Results tracked2135

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

GLM-4.7-Flash: 40.6 (#135), Grok-3 mini: 40.8 (#131)

Coding benchmarks
BenchmarkGLM-4.7-FlashGrok-3 mini
LMArena Coding13831379
Aider Polyglot—49.3%
WeirdML—42.6%

Reasoning GLM-4.7-Flash leads

GLM-4.7-Flash: 20.9 (#229), Grok-3 mini: 13.6 (#334)

Reasoning benchmarks
BenchmarkGLM-4.7-FlashGrok-3 mini
LMArena Hard Prompts13561375
ARC-AGI-2—0.4%
Kagi LLM Benchmark—61.3%
ARC-AGI-1—16.5%
Chess Puzzles0%—
Epoch Capabilities Index—140.35

Math Grok-3 mini leads

GLM-4.7-Flash: 36.1 (#173), Grok-3 mini: 42.1 (#85)

Math benchmarks
BenchmarkGLM-4.7-FlashGrok-3 mini
OTIS Mock AIME 2024-202558.3%77.8%
LMArena Math13551386
Omni-MATH—31.8%
MATH Level 5—90.9%
FrontierMath (Feb 2025 set)—5.9%

Knowledge Grok-3 mini leads

GLM-4.7-Flash: 35.5 (#184), Grok-3 mini: 46.4 (#81)

Knowledge benchmarks
BenchmarkGLM-4.7-FlashGrok-3 mini
GPQA Diamond60.5%76.3%
LMArena Expert13571395
MMLU-Pro—79.9%
Confabulations—10.8%
Vectara Hallucination Rate9.3%—
GPQA (HELM)—67.5%

Multilingual Grok-3 mini leads

GLM-4.7-Flash: 46.5 (#158), Grok-3 mini: 48.1 (#145)

Multilingual benchmarks
BenchmarkGLM-4.7-FlashGrok-3 mini
LMArena Non-English13301352
LMArena Chinese14031387
LMArena French13321357
LMArena German13371349
LMArena Korean12831335
LMArena Russian13321353
LMArena Spanish13501381
LMArena Japanese—1342

Instruction Following Grok-3 mini leads

GLM-4.7-Flash: 70.1 (#167), Grok-3 mini: 78.5 (#9)

Instruction Following benchmarks
BenchmarkGLM-4.7-FlashGrok-3 mini
LMArena Instruction Following13271357
IFEval—95.1%

Long Context Too close to call

GLM-4.7-Flash: 40.9 (#148), Grok-3 mini: 41.0 (#147)

Long Context benchmarks
BenchmarkGLM-4.7-FlashGrok-3 mini
LMArena Longer Query13451372
Fiction.LiveBench—66.7%

Writing & Preference Grok-3 mini leads

GLM-4.7-Flash: 47.4 (#210), Grok-3 mini: 52.5 (#169)

Writing & Preference benchmarks
BenchmarkGLM-4.7-FlashGrok-3 mini
LMArena Text13511370
LMArena Creative Writing12971342
LMArena Multi-Turn13421355
Short-Story Creative Writing—73.5%
EQ-Bench Creative Writing1125—
WildBench—65.1%

Frequently asked questions

Is GLM-4.7-Flash better than Grok-3 mini?

Grok-3 mini is the stronger model overall, scoring 41.2 to 38.8 on the Noometry Index.

Is GLM-4.7-Flash or Grok-3 mini better for coding?

They score almost the same on coding (40.6 vs 40.8); test both on your own repository before choosing.

How many benchmarks do GLM-4.7-Flash and Grok-3 mini share?

18 benchmarks have published results for both models. GLM-4.7-Flash has 21 scored results on Noometry and Grok-3 mini has 35.

Related comparisons

Go deeper