Model comparison

Gemma 2B vs Qwen Max

Qwen Max is the stronger model overall, scoring 34.7 to 29.6 on the Noometry Index.

Last verified . 11 shared benchmarks.

Gemma 2B Google

29.6

Rank #307 Confirmed

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Gemma 2B scores higher in 1 category and Qwen Max in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen Max leads 47.8 to 24.0.
  • Gemma 2B has downloadable open weights; the other is API-only.

Side by side

Gemma 2B and Qwen Max specifications
Gemma 2BQwen Max
ProviderGoogleAlibaba (Qwen)
Noometry Index29.634.7
Released2024-02-212024-04-03
WeightsOpenProprietary
Context window—33K
Max output—8K
Input $ / M tokens—$1.60
Output $ / M tokens—$6.40
Results tracked2323

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen Max leads

Gemma 2B: 29.4 (#305), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkGemma 2BQwen Max
LMArena Coding10101288
Aider Polyglot—21.8%
HumanEval+20.7%—
MBPP+34.1%—

Reasoning Qwen Max leads

Gemma 2B: 18.8 (#275), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkGemma 2BQwen Max
LMArena Hard Prompts9891269
BIG-Bench Hard35.2%—
Epoch Capabilities Index94.2—
HellaSwag71.4%—
PIQA77.3%—
WinoGrande65.4%—

Math Gemma 2B leads

Gemma 2B: 30.0 (#239), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkGemma 2BQwen Max
LMArena Math10091275
OTIS Mock AIME 2024-2025—16.1%
MATH Level 5—67.2%
FrontierMath (Feb 2025 set)—1%
GSM8K17.7%—

Knowledge Not comparable

Gemma 2B: —, Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkGemma 2BQwen Max
GPQA Diamond—56.1%
LMArena Expert—1248
ARC (AI2) Challenge42.1%—
BoolQ69.4%—
MMLU42.3%—
TriviaQA53.2%—

Multilingual Qwen Max leads

Gemma 2B: 23.0 (#294), Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkGemma 2BQwen Max
LMArena Non-English9581263
LMArena Chinese9861254
LMArena Russian9371274
LMArena French—1330
LMArena German—1254
LMArena Japanese—1205
LMArena Korean—1142
LMArena Spanish—1290

Instruction Following Qwen Max leads

Gemma 2B: 48.5 (#302), Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkGemma 2BQwen Max
LMArena Instruction Following9701262

Long Context Qwen Max leads

Gemma 2B: 29.9 (#291), Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkGemma 2BQwen Max
LMArena Longer Query9811288
Fiction.LiveBench—66.7%

Writing & Preference Qwen Max leads

Gemma 2B: 24.0 (#308), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkGemma 2BQwen Max
LMArena Text10021282
LMArena Creative Writing9871248
LMArena Multi-Turn9451277

Frequently asked questions

Is Gemma 2B better than Qwen Max?

Qwen Max is the stronger model overall, scoring 34.7 to 29.6 on the Noometry Index.

Is Gemma 2B or Qwen Max better for coding?

Qwen Max scores higher on coding benchmarks: 30.7 versus 29.4 in the Noometry coding category.

How many benchmarks do Gemma 2B and Qwen Max share?

11 benchmarks have published results for both models. Gemma 2B has 23 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper