Model comparison

Gemma 2 9B vs Qwen Max

Qwen Max is the stronger model overall, scoring 34.7 to 25.9 on the Noometry Index.

Last verified . 20 shared benchmarks.

Gemma 2 9B Google

25.9

Rank #341 Confirmed

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 20 benchmarks with published results for both. Gemma 2 9B scores higher in 0 categories and Qwen Max in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen Max leads 30.3 to 9.7.
  • The biggest single-benchmark swing is MATH Level 5: 21% for Gemma 2 9B and 67.2% for Qwen Max.
  • Gemma 2 9B has downloadable open weights; the other is API-only.

Side by side

Gemma 2 9B and Qwen Max specifications
Gemma 2 9BQwen Max
ProviderGoogleAlibaba (Qwen)
Noometry Index25.934.7
Released2024-06-242024-04-03
WeightsOpenProprietary
Context window—33K
Max output—8K
Input $ / M tokens—$1.60
Output $ / M tokens—$6.40
Results tracked3523

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen Max leads

Gemma 2 9B: 29.4 (#304), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkGemma 2 9BQwen Max
LMArena Coding11731288
Aider Polyglot—21.8%
BigCodeBench Instruct34.7%—
LiveBench Coding22.5%—
BigCodeBench Complete40.6%—

Reasoning Qwen Max leads

Gemma 2 9B: 15.9 (#309), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkGemma 2 9BQwen Max
LMArena Hard Prompts11711269
LiveBench Reasoning15.2%—
LiveBench Data Analysis36.4%—
Epoch Capabilities Index119.83—
LiveBench28.7%—
PIQA83.7%—

Math Qwen Max leads

Gemma 2 9B: 9.9 (#318), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkGemma 2 9BQwen Max
OTIS Mock AIME 2024-20250.6%16.1%
LMArena Math11831275
MATH Level 521%67.2%
LiveBench Math19.8%—
FrontierMath (Feb 2025 set)—1%
GSM8K84.9%—

Knowledge Qwen Max leads

Gemma 2 9B: 9.7 (#305), Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkGemma 2 9BQwen Max
GPQA Diamond27.5%56.1%
LMArena Expert11471248
BoolQ85.7%—
MMLU72.1%—

Multilingual Qwen Max leads

Gemma 2 9B: 36.6 (#238), Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkGemma 2 9BQwen Max
LMArena Non-English11881263
LMArena Chinese11851254
LMArena French11901330
LMArena German11861254
LMArena Japanese11441205
LMArena Korean11371142
LMArena Russian12001274
LMArena Spanish12001290

Instruction Following Qwen Max leads

Gemma 2 9B: 57.6 (#269), Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkGemma 2 9BQwen Max
LMArena Instruction Following11781262
LiveBench Instruction Following52.6%—

Long Context Qwen Max leads

Gemma 2 9B: 36.3 (#233), Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkGemma 2 9BQwen Max
LMArena Longer Query11971288
Fiction.LiveBench—66.7%

Writing & Preference Qwen Max leads

Gemma 2 9B: 32.1 (#281), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkGemma 2 9BQwen Max
LMArena Text12071282
LMArena Creative Writing12061248
LMArena Multi-Turn11931277
EQ-Bench Creative Writing841—
LiveBench Language25.5%—

Frequently asked questions

Is Gemma 2 9B better than Qwen Max?

Qwen Max is the stronger model overall, scoring 34.7 to 25.9 on the Noometry Index.

Is Gemma 2 9B or Qwen Max better for coding?

Qwen Max scores higher on coding benchmarks: 30.7 versus 29.4 in the Noometry coding category.

How many benchmarks do Gemma 2 9B and Qwen Max share?

20 benchmarks have published results for both models. Gemma 2 9B has 35 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper