Model comparison

Gemma 7B vs Qwen-14B

Qwen-14B is the stronger model overall, scoring 31.4 to 30.0 on the Noometry Index.

Last verified . 17 shared benchmarks.

Gemma 7B Google

30.0

Rank #299 Confirmed

Qwen-14B Alibaba (Qwen)

31.4

Rank #275 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Gemma 7B scores higher in 1 category and Qwen-14B in 6 categories; one gap is clear of the uncertainty.

Side by side

Gemma 7B and Qwen-14B specifications
Gemma 7BQwen-14B
ProviderGoogleAlibaba (Qwen)
Noometry Index30.031.4
Released2024-02-212023-09-24
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2718

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Gemma 7B: 30.5 (#294), Qwen-14B: 31.2 (#288)

Coding benchmarks
BenchmarkGemma 7BQwen-14B
LMArena Coding10481071
HumanEval+28.7%—
MBPP+43.4%—

Reasoning Too close to call

Gemma 7B: 19.9 (#249), Qwen-14B: 19.6 (#257)

Reasoning benchmarks
BenchmarkGemma 7BQwen-14B
LMArena Hard Prompts10421027
BIG-Bench Hard55.1%55%
Epoch Capabilities Index111.99113.03
PIQA81.2%79.9%
Adversarial NLI48.7%—
HellaSwag82.2%—
LAMBADA—71.1%
WinoGrande79%—

Math Too close to call

Gemma 7B: 31.2 (#228), Qwen-14B: 31.2 (#227)

Math benchmarks
BenchmarkGemma 7BQwen-14B
LMArena Math10661068
GSM8K46.4%61.3%

Knowledge Not comparable

Gemma 7B: 27.3 (#252), Qwen-14B: —

Knowledge benchmarks
BenchmarkGemma 7BQwen-14B
ARC (AI2) Challenge78.3%84.4%
BoolQ83.2%86.2%
MMLU66.1%66.3%
LMArena Expert1001—
OpenBookQA78.6%—
TriviaQA72.3%—

Multilingual Qwen-14B leads

Gemma 7B: 25.1 (#287), Qwen-14B: 27.5 (#275)

Multilingual benchmarks
BenchmarkGemma 7BQwen-14B
LMArena Non-English9991041
LMArena Chinese10351077
LMArena French1025—
LMArena Russian993—

Instruction Following Too close to call

Gemma 7B: 51.5 (#295), Qwen-14B: 52.4 (#289)

Instruction Following benchmarks
BenchmarkGemma 7BQwen-14B
LMArena Instruction Following10171031

Long Context Too close to call

Gemma 7B: 31.1 (#282), Qwen-14B: 31.3 (#280)

Long Context benchmarks
BenchmarkGemma 7BQwen-14B
LMArena Longer Query10221028

Writing & Preference Too close to call

Gemma 7B: 27.1 (#302), Qwen-14B: 27.6 (#299)

Writing & Preference benchmarks
BenchmarkGemma 7BQwen-14B
LMArena Text10561051
LMArena Creative Writing10241028
LMArena Multi-Turn9631022

Frequently asked questions

Is Gemma 7B better than Qwen-14B?

Qwen-14B is the stronger model overall, scoring 31.4 to 30.0 on the Noometry Index.

Is Gemma 7B or Qwen-14B better for coding?

They score almost the same on coding (30.5 vs 31.2); test both on your own repository before choosing.

How many benchmarks do Gemma 7B and Qwen-14B share?

17 benchmarks have published results for both models. Gemma 7B has 27 scored results on Noometry and Qwen-14B has 18.

Related comparisons

Go deeper