Model comparison

Gemma 3 12B vs Qwen2.5 72B Instruct

Gemma 3 12B and Qwen2.5 72B Instruct score almost the same on the Noometry Index (32.1 vs 31.9), so choose on price, context window or the category you care about most.

Last verified . 17 shared benchmarks.

Gemma 3 12B Google

32.1

Rank #262 Confirmed

Qwen2.5 72B Instruct Alibaba (Qwen)

31.9

Rank #267 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Gemma 3 12B scores higher in 6 categories and Qwen2.5 72B Instruct in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen2.5 72B Instruct leads 22.3 to 15.7.
  • The biggest single-benchmark swing is DTBench: 48.8% for Gemma 3 12B and 62.9% for Qwen2.5 72B Instruct.
  • Gemma 3 12B is cheaper at $0.05 / $0.15 per million input/output tokens, against $1.40 / $5.60 for Qwen2.5 72B Instruct.

Side by side

Gemma 3 12B and Qwen2.5 72B Instruct specifications
Gemma 3 12BQwen2.5 72B Instruct
ProviderGoogleAlibaba (Qwen)
Noometry Index32.131.9
Released2025-03-122024-09
WeightsOpenOpen
Context window131K131K
Max output8K8K
Input $ / M tokens$0.05$1.40
Output $ / M tokens$0.15$5.60
Results tracked2443

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 72B Instruct leads

Gemma 3 12B: 31.7 (#280), Qwen2.5 72B Instruct: 33.2 (#260)

Coding benchmarks
BenchmarkGemma 3 12BQwen2.5 72B Instruct
LMArena Coding12811292
SciCode17.4%—
WeirdML—16%
BigCodeBench Instruct—45.8%
BigCodeBench Complete—55.9%

Agentic & Tool Use Gemma 3 12B leads

Gemma 3 12B: 25.5 (#108), Qwen2.5 72B Instruct: 22.1 (#133)

Agentic & Tool Use benchmarks
BenchmarkGemma 3 12BQwen2.5 72B Instruct
Berkeley Function Calling Leaderboard30.4%—
TheAgentCompany—5.7%
BALROG—16.2%
METR Time Horizons—35.8%

Reasoning Qwen2.5 72B Instruct leads

Gemma 3 12B: 15.7 (#313), Qwen2.5 72B Instruct: 22.3 (#199)

Reasoning benchmarks
BenchmarkGemma 3 12BQwen2.5 72B Instruct
LMArena Hard Prompts13091271
DTBench48.8%62.9%
LMCA4.5%13.4%
Epoch Capabilities Index123.5129
CritPt0%—
Chess Puzzles0%—
BIG-Bench Hard—79.8%
ForecastBench—57.5
HellaSwag—84.8%
PIQA—82.6%
WinoGrande—82.3%

Math Gemma 3 12B leads

Gemma 3 12B: 22.3 (#279), Qwen2.5 72B Instruct: 19.3 (#287)

Math benchmarks
BenchmarkGemma 3 12BQwen2.5 72B Instruct
OTIS Mock AIME 2024-202516.7%8.1%
LMArena Math13071283
Omni-MATH—33%
MATH Level 5—63.2%

Knowledge Too close to call

Gemma 3 12B: 26.5 (#257), Qwen2.5 72B Instruct: 27.0 (#253)

Knowledge benchmarks
BenchmarkGemma 3 12BQwen2.5 72B Instruct
GPQA Diamond39.5%49.1%
LMArena Expert12481245
MMLU-Pro—63.1%
Confabulations—19.1%
Vectara Hallucination Rate4.4%—
GPQA (HELM)—42.6%
ARC (AI2) Challenge—94.5%
MMLU—85.3%
TriviaQA—71.9%

Multimodal Not comparable

Gemma 3 12B: —, Qwen2.5 72B Instruct: —

Multimodal benchmarks
BenchmarkGemma 3 12BQwen2.5 72B Instruct
MindCube46.7%—

Multilingual Gemma 3 12B leads

Gemma 3 12B: 45.7 (#165), Qwen2.5 72B Instruct: 41.0 (#213)

Multilingual benchmarks
BenchmarkGemma 3 12BQwen2.5 72B Instruct
LMArena Non-English13181252
LMArena German13701234
LMArena Russian13351264
LMArena Chinese—1272
LMArena French—1280
LMArena Japanese—1180
LMArena Korean—1188
LMArena Spanish—1256

Instruction Following Gemma 3 12B leads

Gemma 3 12B: 68.6 (#186), Qwen2.5 72B Instruct: 65.5 (#221)

Instruction Following benchmarks
BenchmarkGemma 3 12BQwen2.5 72B Instruct
LMArena Instruction Following12991254
IFEval—80.6%

Long Context Gemma 3 12B leads

Gemma 3 12B: 40.0 (#162), Qwen2.5 72B Instruct: 38.9 (#188)

Long Context benchmarks
BenchmarkGemma 3 12BQwen2.5 72B Instruct
LMArena Longer Query13171282

Writing & Preference Too close to call

Gemma 3 12B: 47.5 (#209), Qwen2.5 72B Instruct: 46.7 (#215)

Writing & Preference benchmarks
BenchmarkGemma 3 12BQwen2.5 72B Instruct
LMArena Text13341269
LMArena Creative Writing13311221
LMArena Multi-Turn13341272
EQ-Bench Creative Writing1126—
WildBench—80.2%

Frequently asked questions

Is Gemma 3 12B better than Qwen2.5 72B Instruct?

Gemma 3 12B and Qwen2.5 72B Instruct score almost the same on the Noometry Index (32.1 vs 31.9), so choose on price, context window or the category you care about most.

Which is cheaper, Gemma 3 12B or Qwen2.5 72B Instruct?

Gemma 3 12B is cheaper. It lists at $0.05 per million input tokens and $0.15 per million output tokens; Qwen2.5 72B Instruct lists at $1.40 and $5.60.

Is Gemma 3 12B or Qwen2.5 72B Instruct better for coding?

Qwen2.5 72B Instruct scores higher on coding benchmarks: 33.2 versus 31.7 in the Noometry coding category.

Which has the bigger context window?

Both accept 131K tokens.

How many benchmarks do Gemma 3 12B and Qwen2.5 72B Instruct share?

17 benchmarks have published results for both models. Gemma 3 12B has 24 scored results on Noometry and Qwen2.5 72B Instruct has 43.

Related comparisons

Go deeper