Model comparison

GLM-4.7-Flash vs Granite 4.2 3b

GLM-4.7-Flash and Granite 4.2 3b score almost the same on the Noometry Index (38.8 vs 39.4), so choose on price, context window or the category you care about most.

Last verified . 11 shared benchmarks.

GLM-4.7-Flash Z.ai (Zhipu)

38.8

Rank #180 Confirmed

Granite 4.2 3b IBM

39.4

Rank #169 Confirmed

Summary

  • They share 11 benchmarks with published results for both. GLM-4.7-Flash scores higher in 5 categories and Granite 4.2 3b in 2 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Granite 4.2 3b leads 26.0 to 20.9.

Side by side

GLM-4.7-Flash and Granite 4.2 3b specifications
GLM-4.7-FlashGranite 4.2 3b
ProviderZ.ai (Zhipu)IBM
Noometry Index38.839.4
Released2026-01-19—
WeightsOpenOpen
Context window200K—
Max output131K—
Input $ / M tokens$0.06—
Output $ / M tokens$0.40—
Results tracked2111

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

GLM-4.7-Flash: 40.6 (#135), Granite 4.2 3b: 40.0 (#151)

Coding benchmarks
BenchmarkGLM-4.7-FlashGranite 4.2 3b
LMArena Coding13831361

Reasoning Granite 4.2 3b leads

GLM-4.7-Flash: 20.9 (#229), Granite 4.2 3b: 26.0 (#138)

Reasoning benchmarks
BenchmarkGLM-4.7-FlashGranite 4.2 3b
LMArena Hard Prompts13561306
Chess Puzzles0%—

Math Not comparable

GLM-4.7-Flash: 36.1 (#173), Granite 4.2 3b: —

Math benchmarks
BenchmarkGLM-4.7-FlashGranite 4.2 3b
OTIS Mock AIME 2024-202558.3%—
LMArena Math1355—

Knowledge Too close to call

GLM-4.7-Flash: 35.5 (#184), Granite 4.2 3b: 36.3 (#171)

Knowledge benchmarks
BenchmarkGLM-4.7-FlashGranite 4.2 3b
LMArena Expert13571315
GPQA Diamond60.5%—
Vectara Hallucination Rate9.3%—

Multilingual GLM-4.7-Flash leads

GLM-4.7-Flash: 46.5 (#158), Granite 4.2 3b: 42.1 (#198)

Multilingual benchmarks
BenchmarkGLM-4.7-FlashGranite 4.2 3b
LMArena Non-English13301268
LMArena Chinese14031269
LMArena Russian13321249
LMArena French1332—
LMArena German1337—
LMArena Korean1283—
LMArena Spanish1350—

Instruction Following GLM-4.7-Flash leads

GLM-4.7-Flash: 70.1 (#167), Granite 4.2 3b: 67.1 (#200)

Instruction Following benchmarks
BenchmarkGLM-4.7-FlashGranite 4.2 3b
LMArena Instruction Following13271273

Long Context GLM-4.7-Flash leads

GLM-4.7-Flash: 40.9 (#148), Granite 4.2 3b: 39.2 (#185)

Long Context benchmarks
BenchmarkGLM-4.7-FlashGranite 4.2 3b
LMArena Longer Query13451291

Writing & Preference Too close to call

GLM-4.7-Flash: 47.4 (#210), Granite 4.2 3b: 47.2 (#212)

Writing & Preference benchmarks
BenchmarkGLM-4.7-FlashGranite 4.2 3b
LMArena Text13511293
LMArena Creative Writing12971205
LMArena Multi-Turn13421290
EQ-Bench Creative Writing1125—

Frequently asked questions

Is GLM-4.7-Flash better than Granite 4.2 3b?

GLM-4.7-Flash and Granite 4.2 3b score almost the same on the Noometry Index (38.8 vs 39.4), so choose on price, context window or the category you care about most.

Is GLM-4.7-Flash or Granite 4.2 3b better for coding?

They score almost the same on coding (40.6 vs 40.0); test both on your own repository before choosing.

How many benchmarks do GLM-4.7-Flash and Granite 4.2 3b share?

11 benchmarks have published results for both models. GLM-4.7-Flash has 21 scored results on Noometry and Granite 4.2 3b has 11.

Related comparisons

Go deeper