Model comparison

Gemma 4 31B IT vs Kimi K2.5

Kimi K2.5 is the stronger model overall, scoring 48.1 to 43.5 on the Noometry Index. Gemma 4 31B IT costs 5.9× less per token, which makes it the better buy when Kimi K2.5's lead doesn't matter for your workload.

Last verified . 31 shared benchmarks.

Gemma 4 31B IT Google

43.5

Rank #90 Confirmed

Kimi K2.5 Moonshot AI

48.1

Rank #57 Confirmed

Summary

  • They share 31 benchmarks with published results for both. Gemma 4 31B IT scores higher in 2 categories and Kimi K2.5 in 7 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Kimi K2.5 leads 53.6 to 37.9.
  • The biggest single-benchmark swing is SimpleQA Verified: 10.4% for Gemma 4 31B IT and 34.3% for Kimi K2.5.
  • Gemma 4 31B IT is cheaper at $0.09 / $0.34 per million input/output tokens, against $0.45 / $2.25 for Kimi K2.5.

Side by side

Gemma 4 31B IT and Kimi K2.5 specifications
Gemma 4 31B ITKimi K2.5
ProviderGoogleMoonshot AI
Noometry Index43.548.1
Released2026-04-022026-01-27
WeightsOpenOpen
Context window262K262K
Max output33K262K
Input $ / M tokens$0.09$0.45
Output $ / M tokens$0.34$2.25
Results tracked3551

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2.5 leads

Gemma 4 31B IT: 42.3 (#108), Kimi K2.5: 48.8 (#53)

Coding benchmarks
BenchmarkGemma 4 31B ITKimi K2.5
LMArena WebDev13661437
SciCode43.4%49%
WeirdML52.3%45.6%
LMArena Coding14591474
ALE-Bench925.5821.65
SWE-bench Verified—73.8%
SWE-bench Verified (bash only)—70.8%
SWE-bench Multilingual—67.3%

Agentic & Tool Use Not comparable

Gemma 4 31B IT: —, Kimi K2.5: 34.2 (#48)

Agentic & Tool Use benchmarks
BenchmarkGemma 4 31B ITKimi K2.5
Terminal-Bench—43.2%
OSWorld—63.3%
Vending-Bench 2—1,198

Reasoning Kimi K2.5 leads

Gemma 4 31B IT: 27.2 (#122), Kimi K2.5: 31.2 (#80)

Reasoning benchmarks
BenchmarkGemma 4 31B ITKimi K2.5
Kagi LLM Benchmark63.5%78.5%
NYT Connections (extended)70.6%69.9%
CritPt1.4%3.1%
Chess Puzzles5%12%
Thematic Generalization53%69.4%
LMArena Hard Prompts14481453
Epoch Capabilities Index142.74148.03
ARC-AGI-2—11.8%
SimpleBench—46.8%
ARC-AGI-1—65.3%
EnigmaEval—3.4%
DTBench82.7%—
LMCA39.3%—
Surface Evolver Bench30.6%—

Math Kimi K2.5 leads

Gemma 4 31B IT: 43.2 (#81), Kimi K2.5: 51.8 (#53)

Math benchmarks
BenchmarkGemma 4 31B ITKimi K2.5
OTIS Mock AIME 2024-202573.3%92.2%
LMArena Math14651470
MathArena Final-Answer Competitions—62.3%
FrontierMath (Feb 2025 set)—27.9%
FrontierMath Tier 4 (v1)—4.2%

Knowledge Kimi K2.5 leads

Gemma 4 31B IT: 37.9 (#151), Kimi K2.5: 53.6 (#56)

Knowledge benchmarks
BenchmarkGemma 4 31B ITKimi K2.5
GPQA Diamond75.8%87.6%
SimpleQA Verified10.4%34.3%
Vectara Hallucination Rate7.4%14.2%
LMArena Expert14651466
Humanity's Last Exam—24.4%

Multimodal Too close to call

Gemma 4 31B IT: 41.6 (#34), Kimi K2.5: 41.1 (#39)

Multimodal benchmarks
BenchmarkGemma 4 31B ITKimi K2.5
LMArena Vision12771269
LMArena Document14251430

Multilingual Too close to call

Gemma 4 31B IT: 53.8 (#57), Kimi K2.5: 53.9 (#53)

Multilingual benchmarks
BenchmarkGemma 4 31B ITKimi K2.5
LMArena Non-English14311433
LMArena Chinese14761495
LMArena French14351454
LMArena Russian14601435
LMArena Spanish14441450
LMArena German—1441
LMArena Japanese—1421
LMArena Korean—1410

Instruction Following Too close to call

Gemma 4 31B IT: 75.5 (#61), Kimi K2.5: 75.3 (#64)

Instruction Following benchmarks
BenchmarkGemma 4 31B ITKimi K2.5
LMArena Instruction Following14331431

Long Context Kimi K2.5 leads

Gemma 4 31B IT: 44.2 (#71), Kimi K2.5: 52.1 (#7)

Long Context benchmarks
BenchmarkGemma 4 31B ITKimi K2.5
LMArena Longer Query14461445
Fiction.LiveBench—86.1%
CL-bench—19.3%
CL-bench Life—13.2%

Writing & Preference Kimi K2.5 leads

Gemma 4 31B IT: 60.5 (#96), Kimi K2.5: 65.1 (#53)

Writing & Preference benchmarks
BenchmarkGemma 4 31B ITKimi K2.5
LMArena Text14431445
LMArena Creative Writing14151423
EQ-Bench Creative Writing13681579
LMArena Multi-Turn14521444
EQ-Bench 41120—

Frequently asked questions

Is Gemma 4 31B IT better than Kimi K2.5?

Kimi K2.5 is the stronger model overall, scoring 48.1 to 43.5 on the Noometry Index. Gemma 4 31B IT costs 5.9× less per token, which makes it the better buy when Kimi K2.5's lead doesn't matter for your workload.

Which is cheaper, Gemma 4 31B IT or Kimi K2.5?

Gemma 4 31B IT is cheaper. It lists at $0.09 per million input tokens and $0.34 per million output tokens; Kimi K2.5 lists at $0.45 and $2.25.

Is Gemma 4 31B IT or Kimi K2.5 better for coding?

Kimi K2.5 scores higher on coding benchmarks: 48.8 versus 42.3 in the Noometry coding category.

Which has the bigger context window?

Both accept 262K tokens.

How many benchmarks do Gemma 4 31B IT and Kimi K2.5 share?

31 benchmarks have published results for both models. Gemma 4 31B IT has 35 scored results on Noometry and Kimi K2.5 has 51.

Related comparisons

Go deeper