Model comparison

Gemma 2 9B vs GLM-4.7-Flash

GLM-4.7-Flash is the stronger model overall, scoring 38.8 to 25.9 on the Noometry Index.

Last verified . 19 shared benchmarks.

Gemma 2 9B Google

25.9

Rank #341 Confirmed

GLM-4.7-Flash Z.ai (Zhipu)

38.8

Rank #180 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Gemma 2 9B scores higher in 0 categories and GLM-4.7-Flash in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where GLM-4.7-Flash leads 36.1 to 9.9.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 0.6% for Gemma 2 9B and 58.3% for GLM-4.7-Flash.

Side by side

Gemma 2 9B and GLM-4.7-Flash specifications
Gemma 2 9BGLM-4.7-Flash
ProviderGoogleZ.ai (Zhipu)
Noometry Index25.938.8
Released2024-06-242026-01-19
WeightsOpenOpen
Context window—200K
Max output—131K
Input $ / M tokens—$0.06
Output $ / M tokens—$0.40
Results tracked3521

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-4.7-Flash leads

Gemma 2 9B: 29.4 (#304), GLM-4.7-Flash: 40.6 (#135)

Coding benchmarks
BenchmarkGemma 2 9BGLM-4.7-Flash
LMArena Coding11731383
BigCodeBench Instruct34.7%—
LiveBench Coding22.5%—
BigCodeBench Complete40.6%—

Reasoning GLM-4.7-Flash leads

Gemma 2 9B: 15.9 (#309), GLM-4.7-Flash: 20.9 (#229)

Reasoning benchmarks
BenchmarkGemma 2 9BGLM-4.7-Flash
LMArena Hard Prompts11711356
Chess Puzzles—0%
LiveBench Reasoning15.2%—
LiveBench Data Analysis36.4%—
Epoch Capabilities Index119.83—
LiveBench28.7%—
PIQA83.7%—

Math GLM-4.7-Flash leads

Gemma 2 9B: 9.9 (#318), GLM-4.7-Flash: 36.1 (#173)

Math benchmarks
BenchmarkGemma 2 9BGLM-4.7-Flash
OTIS Mock AIME 2024-20250.6%58.3%
LMArena Math11831355
LiveBench Math19.8%—
MATH Level 521%—
GSM8K84.9%—

Knowledge GLM-4.7-Flash leads

Gemma 2 9B: 9.7 (#305), GLM-4.7-Flash: 35.5 (#184)

Knowledge benchmarks
BenchmarkGemma 2 9BGLM-4.7-Flash
GPQA Diamond27.5%60.5%
LMArena Expert11471357
Vectara Hallucination Rate—9.3%
BoolQ85.7%—
MMLU72.1%—

Multilingual GLM-4.7-Flash leads

Gemma 2 9B: 36.6 (#238), GLM-4.7-Flash: 46.5 (#158)

Multilingual benchmarks
BenchmarkGemma 2 9BGLM-4.7-Flash
LMArena Non-English11881330
LMArena Chinese11851403
LMArena French11901332
LMArena German11861337
LMArena Korean11371283
LMArena Russian12001332
LMArena Spanish12001350
LMArena Japanese1144—

Instruction Following GLM-4.7-Flash leads

Gemma 2 9B: 57.6 (#269), GLM-4.7-Flash: 70.1 (#167)

Instruction Following benchmarks
BenchmarkGemma 2 9BGLM-4.7-Flash
LMArena Instruction Following11781327
LiveBench Instruction Following52.6%—

Long Context GLM-4.7-Flash leads

Gemma 2 9B: 36.3 (#233), GLM-4.7-Flash: 40.9 (#148)

Long Context benchmarks
BenchmarkGemma 2 9BGLM-4.7-Flash
LMArena Longer Query11971345

Writing & Preference GLM-4.7-Flash leads

Gemma 2 9B: 32.1 (#281), GLM-4.7-Flash: 47.4 (#210)

Writing & Preference benchmarks
BenchmarkGemma 2 9BGLM-4.7-Flash
LMArena Text12071351
LMArena Creative Writing12061297
EQ-Bench Creative Writing8411125
LMArena Multi-Turn11931342
LiveBench Language25.5%—

Frequently asked questions

Is Gemma 2 9B better than GLM-4.7-Flash?

GLM-4.7-Flash is the stronger model overall, scoring 38.8 to 25.9 on the Noometry Index.

Is Gemma 2 9B or GLM-4.7-Flash better for coding?

GLM-4.7-Flash scores higher on coding benchmarks: 40.6 versus 29.4 in the Noometry coding category.

How many benchmarks do Gemma 2 9B and GLM-4.7-Flash share?

19 benchmarks have published results for both models. Gemma 2 9B has 35 scored results on Noometry and GLM-4.7-Flash has 21.

Related comparisons

Go deeper