Model comparison

Gemma 2 9B vs MiMo-V2.5

MiMo-V2.5 is the stronger model overall, scoring 43.4 to 25.9 on the Noometry Index.

Last verified . 17 shared benchmarks.

Gemma 2 9B Google

25.9

Rank #341 Confirmed

MiMo-V2.5 Xiaomi

43.4

Rank #93 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Gemma 2 9B scores higher in 0 categories and MiMo-V2.5 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where MiMo-V2.5 leads 40.8 to 9.7.

Side by side

Gemma 2 9B and MiMo-V2.5 specifications
Gemma 2 9BMiMo-V2.5
ProviderGoogleXiaomi
Noometry Index25.943.4
Released2024-06-242026-04-22
WeightsOpenOpen
Context window—1.05M
Max output—131K
Input $ / M tokens—$0.14
Output $ / M tokens—$0.28
Results tracked3523

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding MiMo-V2.5 leads

Gemma 2 9B: 29.4 (#304), MiMo-V2.5: 43.9 (#81)

Coding benchmarks
BenchmarkGemma 2 9BMiMo-V2.5
LMArena Coding11731469
LMArena WebDev—1438
SciCode—43.1%
BigCodeBench Instruct34.7%—
LiveBench Coding22.5%—
BigCodeBench Complete40.6%—
ALE-Bench—513.95

Reasoning MiMo-V2.5 leads

Gemma 2 9B: 15.9 (#309), MiMo-V2.5: 28.6 (#101)

Reasoning benchmarks
BenchmarkGemma 2 9BMiMo-V2.5
LMArena Hard Prompts11711450
CritPt—3.7%
LiveBench Reasoning15.2%—
LiveBench Data Analysis36.4%—
Epoch Capabilities Index119.83—
LiveBench28.7%—
PIQA83.7%—

Math MiMo-V2.5 leads

Gemma 2 9B: 9.9 (#318), MiMo-V2.5: 36.8 (#163)

Math benchmarks
BenchmarkGemma 2 9BMiMo-V2.5
LMArena Math11831436
OTIS Mock AIME 2024-20250.6%—
ProofBench—16%
LiveBench Math19.8%—
MATH Level 521%—
GSM8K84.9%—

Knowledge MiMo-V2.5 leads

Gemma 2 9B: 9.7 (#305), MiMo-V2.5: 40.8 (#115)

Knowledge benchmarks
BenchmarkGemma 2 9BMiMo-V2.5
LMArena Expert11471460
GPQA Diamond27.5%—
BoolQ85.7%—
MMLU72.1%—

Multimodal Not comparable

Gemma 2 9B: —, MiMo-V2.5: 39.8 (#54)

Multimodal benchmarks
BenchmarkGemma 2 9BMiMo-V2.5
LMArena Vision—1247

Multilingual MiMo-V2.5 leads

Gemma 2 9B: 36.6 (#238), MiMo-V2.5: 51.9 (#99)

Multilingual benchmarks
BenchmarkGemma 2 9BMiMo-V2.5
LMArena Non-English11881404
LMArena Chinese11851468
LMArena French11901447
LMArena German11861421
LMArena Japanese11441306
LMArena Korean11371363
LMArena Russian12001395
LMArena Spanish12001416

Instruction Following MiMo-V2.5 leads

Gemma 2 9B: 57.6 (#269), MiMo-V2.5: 75.5 (#60)

Instruction Following benchmarks
BenchmarkGemma 2 9BMiMo-V2.5
LMArena Instruction Following11781434
LiveBench Instruction Following52.6%—

Long Context MiMo-V2.5 leads

Gemma 2 9B: 36.3 (#233), MiMo-V2.5: 44.2 (#73)

Long Context benchmarks
BenchmarkGemma 2 9BMiMo-V2.5
LMArena Longer Query11971445

Writing & Preference MiMo-V2.5 leads

Gemma 2 9B: 32.1 (#281), MiMo-V2.5: 61.6 (#86)

Writing & Preference benchmarks
BenchmarkGemma 2 9BMiMo-V2.5
LMArena Text12071428
LMArena Creative Writing12061393
LMArena Multi-Turn11931445
EQ-Bench Creative Writing841—
LiveBench Language25.5%—

Frequently asked questions

Is Gemma 2 9B better than MiMo-V2.5?

MiMo-V2.5 is the stronger model overall, scoring 43.4 to 25.9 on the Noometry Index.

Is Gemma 2 9B or MiMo-V2.5 better for coding?

MiMo-V2.5 scores higher on coding benchmarks: 43.9 versus 29.4 in the Noometry coding category.

How many benchmarks do Gemma 2 9B and MiMo-V2.5 share?

17 benchmarks have published results for both models. Gemma 2 9B has 35 scored results on Noometry and MiMo-V2.5 has 23.

Related comparisons

Go deeper