Model comparison

Gemma 2 27B vs Qwen2.5 72B Instruct

Qwen2.5 72B Instruct is the stronger model overall, scoring 31.9 to 29.4 on the Noometry Index. Gemma 2 27B costs 3.8× less per token, which makes it the better buy when Qwen2.5 72B Instruct's lead doesn't matter for your workload.

Last verified . 27 shared benchmarks.

Gemma 2 27B Google

29.4

Rank #312 Confirmed

Qwen2.5 72B Instruct Alibaba (Qwen)

31.9

Rank #267 Confirmed

Summary

  • They share 27 benchmarks with published results for both. Gemma 2 27B scores higher in 1 category and Qwen2.5 72B Instruct in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen2.5 72B Instruct leads 19.3 to 10.7.
  • The biggest single-benchmark swing is MATH Level 5: 27.9% for Gemma 2 27B and 63.2% for Qwen2.5 72B Instruct.
  • Gemma 2 27B is cheaper at $0.65 / $0.65 per million input/output tokens, against $1.40 / $5.60 for Qwen2.5 72B Instruct.
  • Qwen2.5 72B Instruct accepts more context: 131K tokens versus 8K.

Side by side

Gemma 2 27B and Qwen2.5 72B Instruct specifications
Gemma 2 27BQwen2.5 72B Instruct
ProviderGoogleAlibaba (Qwen)
Noometry Index29.431.9
Released2024-06-242024-09
WeightsOpenOpen
Context window8K131K
Max output2K8K
Input $ / M tokens$0.65$1.40
Output $ / M tokens$0.65$5.60
Results tracked3443

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Gemma 2 27B: 34.1 (#246), Qwen2.5 72B Instruct: 33.2 (#260)

Coding benchmarks
BenchmarkGemma 2 27BQwen2.5 72B Instruct
BigCodeBench Instruct42.8%45.8%
LMArena Coding12111292
BigCodeBench Complete52.5%55.9%
WeirdML—16%
LiveBench Coding36%—

Agentic & Tool Use Not comparable

Gemma 2 27B: —, Qwen2.5 72B Instruct: 22.1 (#133)

Agentic & Tool Use benchmarks
BenchmarkGemma 2 27BQwen2.5 72B Instruct
TheAgentCompany—5.7%
BALROG—16.2%
METR Time Horizons—35.8%

Reasoning Qwen2.5 72B Instruct leads

Gemma 2 27B: 15.3 (#315), Qwen2.5 72B Instruct: 22.3 (#199)

Reasoning benchmarks
BenchmarkGemma 2 27BQwen2.5 72B Instruct
LMArena Hard Prompts11981271
DTBench48%62.9%
LMCA7.1%13.4%
Epoch Capabilities Index122.08129
LiveBench Reasoning28.1%—
LiveBench Data Analysis47.9%—
BIG-Bench Hard—79.8%
ForecastBench—57.5
HellaSwag—84.8%
LiveBench38.2%—
PIQA—82.6%
WinoGrande—82.3%

Math Qwen2.5 72B Instruct leads

Gemma 2 27B: 10.7 (#311), Qwen2.5 72B Instruct: 19.3 (#287)

Math benchmarks
BenchmarkGemma 2 27BQwen2.5 72B Instruct
OTIS Mock AIME 2024-20251.4%8.1%
LMArena Math12121283
MATH Level 527.9%63.2%
Omni-MATH—33%
LiveBench Math26.5%—

Knowledge Qwen2.5 72B Instruct leads

Gemma 2 27B: 19.0 (#280), Qwen2.5 72B Instruct: 27.0 (#253)

Knowledge benchmarks
BenchmarkGemma 2 27BQwen2.5 72B Instruct
GPQA Diamond36.5%49.1%
Confabulations27.1%19.1%
LMArena Expert11721245
MMLU75.7%85.3%
MMLU-Pro—63.1%
GPQA (HELM)—42.6%
ARC (AI2) Challenge—94.5%
TriviaQA—71.9%

Multilingual Qwen2.5 72B Instruct leads

Gemma 2 27B: 38.6 (#226), Qwen2.5 72B Instruct: 41.0 (#213)

Multilingual benchmarks
BenchmarkGemma 2 27BQwen2.5 72B Instruct
LMArena Non-English12171252
LMArena Chinese12211272
LMArena French12471280
LMArena German12091234
LMArena Japanese11751180
LMArena Korean11741188
LMArena Russian12341264
LMArena Spanish12281256

Instruction Following Qwen2.5 72B Instruct leads

Gemma 2 27B: 60.5 (#249), Qwen2.5 72B Instruct: 65.5 (#221)

Instruction Following benchmarks
BenchmarkGemma 2 27BQwen2.5 72B Instruct
LMArena Instruction Following12061254
LiveBench Instruction Following58.1%—
IFEval—80.6%

Long Context Qwen2.5 72B Instruct leads

Gemma 2 27B: 37.3 (#218), Qwen2.5 72B Instruct: 38.9 (#188)

Long Context benchmarks
BenchmarkGemma 2 27BQwen2.5 72B Instruct
LMArena Longer Query12311282

Writing & Preference Qwen2.5 72B Instruct leads

Gemma 2 27B: 44.2 (#225), Qwen2.5 72B Instruct: 46.7 (#215)

Writing & Preference benchmarks
BenchmarkGemma 2 27BQwen2.5 72B Instruct
LMArena Text12311269
LMArena Creative Writing12411221
LMArena Multi-Turn12241272
WildBench—80.2%
LiveBench Language32.6%—

Frequently asked questions

Is Gemma 2 27B better than Qwen2.5 72B Instruct?

Qwen2.5 72B Instruct is the stronger model overall, scoring 31.9 to 29.4 on the Noometry Index. Gemma 2 27B costs 3.8× less per token, which makes it the better buy when Qwen2.5 72B Instruct's lead doesn't matter for your workload.

Which is cheaper, Gemma 2 27B or Qwen2.5 72B Instruct?

Gemma 2 27B is cheaper. It lists at $0.65 per million input tokens and $0.65 per million output tokens; Qwen2.5 72B Instruct lists at $1.40 and $5.60.

Is Gemma 2 27B or Qwen2.5 72B Instruct better for coding?

They score almost the same on coding (34.1 vs 33.2); test both on your own repository before choosing.

Which has the bigger context window?

Qwen2.5 72B Instruct does, with 131K tokens against 8K.

How many benchmarks do Gemma 2 27B and Qwen2.5 72B Instruct share?

27 benchmarks have published results for both models. Gemma 2 27B has 34 scored results on Noometry and Qwen2.5 72B Instruct has 43.

Related comparisons

Go deeper