Model comparison

Gemma 1.1 2b IT vs Qwen Max

Qwen Max is the stronger model overall, scoring 34.7 to 29.3 on the Noometry Index.

Last verified . 14 shared benchmarks.

Gemma 1.1 2b IT Google

29.3

Rank #313 Confirmed

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Gemma 1.1 2b IT scores higher in 1 category and Qwen Max in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen Max leads 47.8 to 25.1.
  • Gemma 1.1 2b IT has downloadable open weights; the other is API-only.

Side by side

Gemma 1.1 2b IT and Qwen Max specifications
Gemma 1.1 2b ITQwen Max
ProviderGoogleAlibaba (Qwen)
Noometry Index29.334.7
Released—2024-04-03
WeightsOpenProprietary
Context window—33K
Max output—8K
Input $ / M tokens—$1.60
Output $ / M tokens—$6.40
Results tracked1623

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Gemma 1.1 2b IT: 30.1 (#299), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkGemma 1.1 2b ITQwen Max
LMArena Coding10341288
Aider Polyglot—21.8%
HumanEval+17.7%—
MBPP+23.3%—

Reasoning Qwen Max leads

Gemma 1.1 2b IT: 19.1 (#270), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkGemma 1.1 2b ITQwen Max
LMArena Hard Prompts10051269

Math Gemma 1.1 2b IT leads

Gemma 1.1 2b IT: 30.8 (#232), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkGemma 1.1 2b ITQwen Max
LMArena Math10471275
OTIS Mock AIME 2024-2025—16.1%
MATH Level 5—67.2%
FrontierMath (Feb 2025 set)—1%

Knowledge Qwen Max leads

Gemma 1.1 2b IT: 26.5 (#258), Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkGemma 1.1 2b ITQwen Max
LMArena Expert9701248
GPQA Diamond—56.1%

Multilingual Qwen Max leads

Gemma 1.1 2b IT: 24.6 (#289), Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkGemma 1.1 2b ITQwen Max
LMArena Non-English9881263
LMArena Chinese10121254
LMArena German9441254
LMArena Korean8991142
LMArena Russian9901274
LMArena French—1330
LMArena Japanese—1205
LMArena Spanish—1290

Instruction Following Qwen Max leads

Gemma 1.1 2b IT: 49.9 (#299), Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkGemma 1.1 2b ITQwen Max
LMArena Instruction Following9921262

Long Context Qwen Max leads

Gemma 1.1 2b IT: 30.6 (#286), Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkGemma 1.1 2b ITQwen Max
LMArena Longer Query10031288
Fiction.LiveBench—66.7%

Writing & Preference Qwen Max leads

Gemma 1.1 2b IT: 25.1 (#306), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkGemma 1.1 2b ITQwen Max
LMArena Text10221282
LMArena Creative Writing9981248
LMArena Multi-Turn9591277

Frequently asked questions

Is Gemma 1.1 2b IT better than Qwen Max?

Qwen Max is the stronger model overall, scoring 34.7 to 29.3 on the Noometry Index.

Is Gemma 1.1 2b IT or Qwen Max better for coding?

They score almost the same on coding (30.1 vs 30.7); test both on your own repository before choosing.

How many benchmarks do Gemma 1.1 2b IT and Qwen Max share?

14 benchmarks have published results for both models. Gemma 1.1 2b IT has 16 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper