Model comparison

Gemma 2 9B vs GLM-5V-Turbo

GLM-5V-Turbo is the stronger model overall, scoring 43.8 to 25.9 on the Noometry Index.

Last verified . 16 shared benchmarks.

Gemma 2 9B Google

25.9

Rank #341 Confirmed

GLM-5V-Turbo Z.ai (Zhipu)

43.8

Rank #84 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Gemma 2 9B scores higher in 0 categories and GLM-5V-Turbo in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where GLM-5V-Turbo leads 40.6 to 9.7.
  • Gemma 2 9B has downloadable open weights; the other is API-only.

Side by side

Gemma 2 9B and GLM-5V-Turbo specifications
Gemma 2 9BGLM-5V-Turbo
ProviderGoogleZ.ai (Zhipu)
Noometry Index25.943.8
Released2024-06-242026-04-01
WeightsOpenProprietary
Context window—200K
Max output—131K
Input $ / M tokens—$1.20
Output $ / M tokens—$4
Results tracked3519

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GLM-5V-Turbo leads

Gemma 2 9B: 29.4 (#304), GLM-5V-Turbo: 42.1 (#111)

Coding benchmarks
BenchmarkGemma 2 9BGLM-5V-Turbo
LMArena Coding11731466
LMArena WebDev—1401
BigCodeBench Instruct34.7%—
LiveBench Coding22.5%—
BigCodeBench Complete40.6%—

Reasoning GLM-5V-Turbo leads

Gemma 2 9B: 15.9 (#309), GLM-5V-Turbo: 29.7 (#89)

Reasoning benchmarks
BenchmarkGemma 2 9BGLM-5V-Turbo
LMArena Hard Prompts11711443
LiveBench Reasoning15.2%—
LiveBench Data Analysis36.4%—
Epoch Capabilities Index119.83—
LiveBench28.7%—
PIQA83.7%—

Math GLM-5V-Turbo leads

Gemma 2 9B: 9.9 (#318), GLM-5V-Turbo: 39.4 (#106)

Math benchmarks
BenchmarkGemma 2 9BGLM-5V-Turbo
LMArena Math11831441
OTIS Mock AIME 2024-20250.6%—
LiveBench Math19.8%—
MATH Level 521%—
GSM8K84.9%—

Knowledge GLM-5V-Turbo leads

Gemma 2 9B: 9.7 (#305), GLM-5V-Turbo: 40.6 (#117)

Knowledge benchmarks
BenchmarkGemma 2 9BGLM-5V-Turbo
LMArena Expert11471452
GPQA Diamond27.5%—
BoolQ85.7%—
MMLU72.1%—

Multimodal Not comparable

Gemma 2 9B: —, GLM-5V-Turbo: 40.9 (#42)

Multimodal benchmarks
BenchmarkGemma 2 9BGLM-5V-Turbo
LMArena Vision—1264
LMArena Document—1416

Multilingual GLM-5V-Turbo leads

Gemma 2 9B: 36.6 (#238), GLM-5V-Turbo: 53.0 (#73)

Multilingual benchmarks
BenchmarkGemma 2 9BGLM-5V-Turbo
LMArena Non-English11881420
LMArena Chinese11851488
LMArena French11901444
LMArena German11861423
LMArena Korean11371396
LMArena Russian12001431
LMArena Spanish12001450
LMArena Japanese1144—

Instruction Following GLM-5V-Turbo leads

Gemma 2 9B: 57.6 (#269), GLM-5V-Turbo: 75.0 (#80)

Instruction Following benchmarks
BenchmarkGemma 2 9BGLM-5V-Turbo
LMArena Instruction Following11781423
LiveBench Instruction Following52.6%—

Long Context GLM-5V-Turbo leads

Gemma 2 9B: 36.3 (#233), GLM-5V-Turbo: 44.0 (#80)

Long Context benchmarks
BenchmarkGemma 2 9BGLM-5V-Turbo
LMArena Longer Query11971438

Writing & Preference GLM-5V-Turbo leads

Gemma 2 9B: 32.1 (#281), GLM-5V-Turbo: 62.5 (#73)

Writing & Preference benchmarks
BenchmarkGemma 2 9BGLM-5V-Turbo
LMArena Text12071437
LMArena Creative Writing12061416
LMArena Multi-Turn11931432
EQ-Bench Creative Writing841—
LiveBench Language25.5%—

Frequently asked questions

Is Gemma 2 9B better than GLM-5V-Turbo?

GLM-5V-Turbo is the stronger model overall, scoring 43.8 to 25.9 on the Noometry Index.

Is Gemma 2 9B or GLM-5V-Turbo better for coding?

GLM-5V-Turbo scores higher on coding benchmarks: 42.1 versus 29.4 in the Noometry coding category.

How many benchmarks do Gemma 2 9B and GLM-5V-Turbo share?

16 benchmarks have published results for both models. Gemma 2 9B has 35 scored results on Noometry and GLM-5V-Turbo has 19.

Related comparisons

Go deeper