Model comparison

Gemma 2 27B vs Qwen2.5 7B Instruct

Gemma 2 27B and Qwen2.5 7B Instruct score almost the same on the Noometry Index (29.4 vs 29.0), so choose on price, context window or the category you care about most.

Last verified . 8 shared benchmarks.

Gemma 2 27B Google

29.4

Rank #312 Confirmed

Qwen2.5 7B Instruct Alibaba (Qwen)

29.0

Rank #320 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Gemma 2 27B scores higher in 2 categories and Qwen2.5 7B Instruct in 4 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen2.5 7B Instruct leads 48.8 to 44.2.
  • The biggest single-benchmark swing is BigCodeBench Complete: 52.5% for Gemma 2 27B and 46.1% for Qwen2.5 7B Instruct.
  • Qwen2.5 7B Instruct is cheaper at $0.17 / $0.70 per million input/output tokens, against $0.65 / $0.65 for Gemma 2 27B.
  • Qwen2.5 7B Instruct accepts more context: 131K tokens versus 8K.

Side by side

Gemma 2 27B and Qwen2.5 7B Instruct specifications
Gemma 2 27BQwen2.5 7B Instruct
ProviderGoogleAlibaba (Qwen)
Noometry Index29.429.0
Released2024-06-242024-09
WeightsOpenOpen
Context window8K131K
Max output2K8K
Input $ / M tokens$0.65$0.17
Output $ / M tokens$0.65$0.70
Results tracked3415

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 7B Instruct leads

Gemma 2 27B: 34.1 (#246), Qwen2.5 7B Instruct: 36.5 (#208)

Coding benchmarks
BenchmarkGemma 2 27BQwen2.5 7B Instruct
BigCodeBench Instruct42.8%37.6%
BigCodeBench Complete52.5%46.1%
LiveBench Coding36%—
LMArena Coding1211—

Agentic & Tool Use Not comparable

Gemma 2 27B: —, Qwen2.5 7B Instruct: 23.8 (#124)

Agentic & Tool Use benchmarks
BenchmarkGemma 2 27BQwen2.5 7B Instruct
BALROG—7.8%

Reasoning Too close to call

Gemma 2 27B: 15.3 (#315), Qwen2.5 7B Instruct: 14.8 (#322)

Reasoning benchmarks
BenchmarkGemma 2 27BQwen2.5 7B Instruct
DTBench48%47.7%
LMCA7.1%6.4%
Epoch Capabilities Index122.08118.51
Chess Puzzles—0%
LiveBench Reasoning28.1%—
LMArena Hard Prompts1198—
LiveBench Data Analysis47.9%—
LiveBench38.2%—

Math Qwen2.5 7B Instruct leads

Gemma 2 27B: 10.7 (#311), Qwen2.5 7B Instruct: 12.6 (#306)

Math benchmarks
BenchmarkGemma 2 27BQwen2.5 7B Instruct
OTIS Mock AIME 2024-20251.4%2.5%
Omni-MATH—29.4%
LiveBench Math26.5%—
LMArena Math1212—
MATH Level 527.9%—

Knowledge Gemma 2 27B leads

Gemma 2 27B: 19.0 (#280), Qwen2.5 7B Instruct: 17.0 (#286)

Knowledge benchmarks
BenchmarkGemma 2 27BQwen2.5 7B Instruct
GPQA Diamond36.5%35.5%
MMLU75.7%72.9%
MMLU-Pro—53.9%
Confabulations27.1%—
GPQA (HELM)—34.1%
LMArena Expert1172—

Multilingual Not comparable

Gemma 2 27B: 38.6 (#226), Qwen2.5 7B Instruct: —

Multilingual benchmarks
BenchmarkGemma 2 27BQwen2.5 7B Instruct
LMArena Non-English1217—
LMArena Chinese1221—
LMArena French1247—
LMArena German1209—
LMArena Japanese1175—
LMArena Korean1174—
LMArena Russian1234—
LMArena Spanish1228—

Instruction Following Qwen2.5 7B Instruct leads

Gemma 2 27B: 60.5 (#249), Qwen2.5 7B Instruct: 63.2 (#231)

Instruction Following benchmarks
BenchmarkGemma 2 27BQwen2.5 7B Instruct
LiveBench Instruction Following58.1%—
IFEval—74.1%
LMArena Instruction Following1206—

Long Context Not comparable

Gemma 2 27B: 37.3 (#218), Qwen2.5 7B Instruct: —

Long Context benchmarks
BenchmarkGemma 2 27BQwen2.5 7B Instruct
LMArena Longer Query1231—

Writing & Preference Qwen2.5 7B Instruct leads

Gemma 2 27B: 44.2 (#225), Qwen2.5 7B Instruct: 48.8 (#195)

Writing & Preference benchmarks
BenchmarkGemma 2 27BQwen2.5 7B Instruct
LMArena Text1231—
LMArena Creative Writing1241—
WildBench—73.1%
LMArena Multi-Turn1224—
LiveBench Language32.6%—

Frequently asked questions

Is Gemma 2 27B better than Qwen2.5 7B Instruct?

Gemma 2 27B and Qwen2.5 7B Instruct score almost the same on the Noometry Index (29.4 vs 29.0), so choose on price, context window or the category you care about most.

Which is cheaper, Gemma 2 27B or Qwen2.5 7B Instruct?

Qwen2.5 7B Instruct is cheaper. It lists at $0.17 per million input tokens and $0.70 per million output tokens; Gemma 2 27B lists at $0.65 and $0.65.

Is Gemma 2 27B or Qwen2.5 7B Instruct better for coding?

Qwen2.5 7B Instruct scores higher on coding benchmarks: 36.5 versus 34.1 in the Noometry coding category.

Which has the bigger context window?

Qwen2.5 7B Instruct does, with 131K tokens against 8K.

How many benchmarks do Gemma 2 27B and Qwen2.5 7B Instruct share?

8 benchmarks have published results for both models. Gemma 2 27B has 34 scored results on Noometry and Qwen2.5 7B Instruct has 15.

Related comparisons

Go deeper