Model comparison

Gemma 4 31B IT vs GLM-4.5V

Gemma 4 31B IT is the stronger model overall, scoring 43.5 to 39.8 on the Noometry Index.

Last verified . 15 shared benchmarks.

Gemma 4 31B IT Google

43.5

Rank #90 Confirmed

GLM-4.5V Z.ai (Zhipu)

39.8

Rank #158 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Gemma 4 31B IT scores higher in 8 categories and GLM-4.5V in 1 category; 7 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Gemma 4 31B IT leads 53.8 to 44.6.
  • Gemma 4 31B IT is cheaper at $0.09 / $0.34 per million input/output tokens, against $0.60 / $1.80 for GLM-4.5V.
  • Gemma 4 31B IT accepts more context: 262K tokens versus 64K.

Side by side

Gemma 4 31B IT and GLM-4.5V specifications
Gemma 4 31B ITGLM-4.5V
ProviderGoogleZ.ai (Zhipu)
Noometry Index43.539.8
Released2026-04-022025-08-11
WeightsOpenOpen
Context window262K64K
Max output33K16K
Input $ / M tokens$0.09$0.60
Output $ / M tokens$0.34$1.80
Results tracked3515

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 4 31B IT leads

Gemma 4 31B IT: 42.3 (#108), GLM-4.5V: 39.5 (#155)

Coding benchmarks
BenchmarkGemma 4 31B ITGLM-4.5V
LMArena Coding14591347
LMArena WebDev1366—
SciCode43.4%—
WeirdML52.3%—
ALE-Bench925.5—

Reasoning Too close to call

Gemma 4 31B IT: 27.2 (#122), GLM-4.5V: 27.4 (#119)

Reasoning benchmarks
BenchmarkGemma 4 31B ITGLM-4.5V
Kagi LLM Benchmark63.5%59.8%
LMArena Hard Prompts14481334
NYT Connections (extended)70.6%—
CritPt1.4%—
Chess Puzzles5%—
Thematic Generalization53%—
DTBench82.7%—
LMCA39.3%—
Surface Evolver Bench30.6%—
Epoch Capabilities Index142.74—

Math Gemma 4 31B IT leads

Gemma 4 31B IT: 43.2 (#81), GLM-4.5V: 37.4 (#159)

Math benchmarks
BenchmarkGemma 4 31B ITGLM-4.5V
LMArena Math14651354
OTIS Mock AIME 2024-202573.3%—

Knowledge Too close to call

Gemma 4 31B IT: 37.9 (#151), GLM-4.5V: 37.5 (#156)

Knowledge benchmarks
BenchmarkGemma 4 31B ITGLM-4.5V
LMArena Expert14651353
GPQA Diamond75.8%—
SimpleQA Verified10.4%—
Vectara Hallucination Rate7.4%—

Multimodal Gemma 4 31B IT leads

Gemma 4 31B IT: 41.6 (#34), GLM-4.5V: 34.3 (#92)

Multimodal benchmarks
BenchmarkGemma 4 31B ITGLM-4.5V
LMArena Vision12771154
LMArena Document1425—

Multilingual Gemma 4 31B IT leads

Gemma 4 31B IT: 53.8 (#57), GLM-4.5V: 44.6 (#177)

Multilingual benchmarks
BenchmarkGemma 4 31B ITGLM-4.5V
LMArena Non-English14311303
LMArena Chinese14761337
LMArena Russian14601298
LMArena Spanish14441336
LMArena French1435—

Instruction Following Gemma 4 31B IT leads

Gemma 4 31B IT: 75.5 (#61), GLM-4.5V: 69.2 (#175)

Instruction Following benchmarks
BenchmarkGemma 4 31B ITGLM-4.5V
LMArena Instruction Following14331311

Long Context Gemma 4 31B IT leads

Gemma 4 31B IT: 44.2 (#71), GLM-4.5V: 39.6 (#171)

Long Context benchmarks
BenchmarkGemma 4 31B ITGLM-4.5V
LMArena Longer Query14461304

Writing & Preference Gemma 4 31B IT leads

Gemma 4 31B IT: 60.5 (#96), GLM-4.5V: 52.5 (#170)

Writing & Preference benchmarks
BenchmarkGemma 4 31B ITGLM-4.5V
LMArena Text14431333
LMArena Creative Writing14151295
LMArena Multi-Turn14521332
EQ-Bench Creative Writing1368—
EQ-Bench 41120—

Frequently asked questions

Is Gemma 4 31B IT better than GLM-4.5V?

Gemma 4 31B IT is the stronger model overall, scoring 43.5 to 39.8 on the Noometry Index.

Which is cheaper, Gemma 4 31B IT or GLM-4.5V?

Gemma 4 31B IT is cheaper. It lists at $0.09 per million input tokens and $0.34 per million output tokens; GLM-4.5V lists at $0.60 and $1.80.

Is Gemma 4 31B IT or GLM-4.5V better for coding?

Gemma 4 31B IT scores higher on coding benchmarks: 42.3 versus 39.5 in the Noometry coding category.

Which has the bigger context window?

Gemma 4 31B IT does, with 262K tokens against 64K.

How many benchmarks do Gemma 4 31B IT and GLM-4.5V share?

15 benchmarks have published results for both models. Gemma 4 31B IT has 35 scored results on Noometry and GLM-4.5V has 15.

Related comparisons

Go deeper