Model comparison

Gemini 1.5 Pro (May 2024) vs Kimi K2.5 Instant

Kimi K2.5 Instant is the stronger model overall, scoring 43.6 to 32.1 on the Noometry Index.

Last verified . 17 shared benchmarks.

Gemini 1.5 Pro (May 2024) Google

32.1

Rank #261 Confirmed

Kimi K2.5 Instant Moonshot AI

43.6

Rank #89 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Gemini 1.5 Pro (May 2024) scores higher in 0 categories and Kimi K2.5 Instant in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Kimi K2.5 Instant leads 29.7 to 12.3.
  • Kimi K2.5 Instant has downloadable open weights; the other is API-only.

Side by side

Gemini 1.5 Pro (May 2024) and Kimi K2.5 Instant specifications
Gemini 1.5 Pro (May 2024)Kimi K2.5 Instant
ProviderGoogleMoonshot AI
Noometry Index32.143.6
Released2024-02-15—
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4518

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2.5 Instant leads

Gemini 1.5 Pro (May 2024): 34.2 (#241), Kimi K2.5 Instant: 42.6 (#97)

Coding benchmarks
BenchmarkGemini 1.5 Pro (May 2024)Kimi K2.5 Instant
LMArena Coding12941484
LMArena WebDev—1404
WeirdML22.2%—
BigCodeBench Instruct43.8%—
BigCodeBench Complete57.5%—
CadEval34%—
HumanEval+79.3%—
MBPP+74.6%—

Agentic & Tool Use Not comparable

Gemini 1.5 Pro (May 2024): 17.9 (#145), Kimi K2.5 Instant: —

Agentic & Tool Use benchmarks
BenchmarkGemini 1.5 Pro (May 2024)Kimi K2.5 Instant
TheAgentCompany3.4%—
Cybench7.5%—
BALROG21%—

Reasoning Kimi K2.5 Instant leads

Gemini 1.5 Pro (May 2024): 12.3 (#338), Kimi K2.5 Instant: 29.7 (#90)

Reasoning benchmarks
BenchmarkGemini 1.5 Pro (May 2024)Kimi K2.5 Instant
LMArena Hard Prompts12961443
ARC-AGI-20.8%—
SimpleBench27.1%—
DTBench59%—
BIG-Bench Hard89.2%—
Epoch Capabilities Index131.73—
ForecastBench58.4—

Math Kimi K2.5 Instant leads

Gemini 1.5 Pro (May 2024): 25.8 (#266), Kimi K2.5 Instant: 39.4 (#105)

Math benchmarks
BenchmarkGemini 1.5 Pro (May 2024)Kimi K2.5 Instant
LMArena Math13151442
OTIS Mock AIME 2024-202523.1%—
Omni-MATH36.4%—
MATH Level 570.4%—

Knowledge Kimi K2.5 Instant leads

Gemini 1.5 Pro (May 2024): 29.4 (#239), Kimi K2.5 Instant: 40.2 (#123)

Knowledge benchmarks
BenchmarkGemini 1.5 Pro (May 2024)Kimi K2.5 Instant
LMArena Expert12791440
GPQA Diamond57.2%—
Humanity's Last Exam4.6%—
MMLU-Pro73.7%—
Confabulations13.5%—
GPQA (HELM)53.4%—
MMLU86.9%—

Multimodal Kimi K2.5 Instant leads

Gemini 1.5 Pro (May 2024): 36.8 (#77), Kimi K2.5 Instant: 40.2 (#50)

Multimodal benchmarks
BenchmarkGemini 1.5 Pro (May 2024)Kimi K2.5 Instant
LMArena Vision11611254
Video-MME75%—

Multilingual Kimi K2.5 Instant leads

Gemini 1.5 Pro (May 2024): 45.3 (#174), Kimi K2.5 Instant: 52.0 (#94)

Multilingual benchmarks
BenchmarkGemini 1.5 Pro (May 2024)Kimi K2.5 Instant
LMArena Non-English13121406
LMArena Chinese13311449
LMArena French13021403
LMArena German12861413
LMArena Korean12981378
LMArena Russian13201404
LMArena Spanish13111447
LMArena Japanese1292—

Instruction Following Kimi K2.5 Instant leads

Gemini 1.5 Pro (May 2024): 68.6 (#185), Kimi K2.5 Instant: 75.3 (#65)

Instruction Following benchmarks
BenchmarkGemini 1.5 Pro (May 2024)Kimi K2.5 Instant
LMArena Instruction Following12971430
IFEval83.7%—

Long Context Kimi K2.5 Instant leads

Gemini 1.5 Pro (May 2024): 39.8 (#169), Kimi K2.5 Instant: 43.9 (#83)

Long Context benchmarks
BenchmarkGemini 1.5 Pro (May 2024)Kimi K2.5 Instant
LMArena Longer Query13081435

Writing & Preference Kimi K2.5 Instant leads

Gemini 1.5 Pro (May 2024): 52.4 (#172), Kimi K2.5 Instant: 60.6 (#95)

Writing & Preference benchmarks
BenchmarkGemini 1.5 Pro (May 2024)Kimi K2.5 Instant
LMArena Text13191420
LMArena Creative Writing13331381
LMArena Multi-Turn12961427
WildBench81.3%—

Frequently asked questions

Is Gemini 1.5 Pro (May 2024) better than Kimi K2.5 Instant?

Kimi K2.5 Instant is the stronger model overall, scoring 43.6 to 32.1 on the Noometry Index.

Is Gemini 1.5 Pro (May 2024) or Kimi K2.5 Instant better for coding?

Kimi K2.5 Instant scores higher on coding benchmarks: 42.6 versus 34.2 in the Noometry coding category.

How many benchmarks do Gemini 1.5 Pro (May 2024) and Kimi K2.5 Instant share?

17 benchmarks have published results for both models. Gemini 1.5 Pro (May 2024) has 45 scored results on Noometry and Kimi K2.5 Instant has 18.

Related comparisons

Go deeper