Model comparison

Gemma 7B vs Kimi K2.5 Instant

Kimi K2.5 Instant is the stronger model overall, scoring 43.6 to 30.0 on the Noometry Index.

Last verified . 13 shared benchmarks.

Gemma 7B Google

30.0

Rank #299 Confirmed

Kimi K2.5 Instant Moonshot AI

43.6

Rank #89 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Gemma 7B scores higher in 0 categories and Kimi K2.5 Instant in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Kimi K2.5 Instant leads 60.6 to 27.1.

Side by side

Gemma 7B and Kimi K2.5 Instant specifications
Gemma 7BKimi K2.5 Instant
ProviderGoogleMoonshot AI
Noometry Index30.043.6
Released2024-02-21—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2718

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2.5 Instant leads

Gemma 7B: 30.5 (#294), Kimi K2.5 Instant: 42.6 (#97)

Coding benchmarks
BenchmarkGemma 7BKimi K2.5 Instant
LMArena Coding10481484
LMArena WebDev—1404
HumanEval+28.7%—
MBPP+43.4%—

Reasoning Kimi K2.5 Instant leads

Gemma 7B: 19.9 (#249), Kimi K2.5 Instant: 29.7 (#90)

Reasoning benchmarks
BenchmarkGemma 7BKimi K2.5 Instant
LMArena Hard Prompts10421443
Adversarial NLI48.7%—
BIG-Bench Hard55.1%—
Epoch Capabilities Index111.99—
HellaSwag82.2%—
PIQA81.2%—
WinoGrande79%—

Math Kimi K2.5 Instant leads

Gemma 7B: 31.2 (#228), Kimi K2.5 Instant: 39.4 (#105)

Math benchmarks
BenchmarkGemma 7BKimi K2.5 Instant
LMArena Math10661442
GSM8K46.4%—

Knowledge Kimi K2.5 Instant leads

Gemma 7B: 27.3 (#252), Kimi K2.5 Instant: 40.2 (#123)

Knowledge benchmarks
BenchmarkGemma 7BKimi K2.5 Instant
LMArena Expert10011440
ARC (AI2) Challenge78.3%—
BoolQ83.2%—
MMLU66.1%—
OpenBookQA78.6%—
TriviaQA72.3%—

Multimodal Not comparable

Gemma 7B: —, Kimi K2.5 Instant: 40.2 (#50)

Multimodal benchmarks
BenchmarkGemma 7BKimi K2.5 Instant
LMArena Vision—1254

Multilingual Kimi K2.5 Instant leads

Gemma 7B: 25.1 (#287), Kimi K2.5 Instant: 52.0 (#94)

Multilingual benchmarks
BenchmarkGemma 7BKimi K2.5 Instant
LMArena Non-English9991406
LMArena Chinese10351449
LMArena French10251403
LMArena Russian9931404
LMArena German—1413
LMArena Korean—1378
LMArena Spanish—1447

Instruction Following Kimi K2.5 Instant leads

Gemma 7B: 51.5 (#295), Kimi K2.5 Instant: 75.3 (#65)

Instruction Following benchmarks
BenchmarkGemma 7BKimi K2.5 Instant
LMArena Instruction Following10171430

Long Context Kimi K2.5 Instant leads

Gemma 7B: 31.1 (#282), Kimi K2.5 Instant: 43.9 (#83)

Long Context benchmarks
BenchmarkGemma 7BKimi K2.5 Instant
LMArena Longer Query10221435

Writing & Preference Kimi K2.5 Instant leads

Gemma 7B: 27.1 (#302), Kimi K2.5 Instant: 60.6 (#95)

Writing & Preference benchmarks
BenchmarkGemma 7BKimi K2.5 Instant
LMArena Text10561420
LMArena Creative Writing10241381
LMArena Multi-Turn9631427

Frequently asked questions

Is Gemma 7B better than Kimi K2.5 Instant?

Kimi K2.5 Instant is the stronger model overall, scoring 43.6 to 30.0 on the Noometry Index.

Is Gemma 7B or Kimi K2.5 Instant better for coding?

Kimi K2.5 Instant scores higher on coding benchmarks: 42.6 versus 30.5 in the Noometry coding category.

How many benchmarks do Gemma 7B and Kimi K2.5 Instant share?

13 benchmarks have published results for both models. Gemma 7B has 27 scored results on Noometry and Kimi K2.5 Instant has 18.

Related comparisons

Go deeper