Model comparison

Gemini 2.5 Flash-Lite vs Kimi K2.5

Kimi K2.5 is the stronger model overall, scoring 48.1 to 37.0 on the Noometry Index. Gemini 2.5 Flash-Lite costs 5.1× less per token, which makes it the better buy when Kimi K2.5's lead doesn't matter for your workload.

Last verified . 24 shared benchmarks.

Gemini 2.5 Flash-Lite Google

37.0

Rank #211 Confirmed

Kimi K2.5 Moonshot AI

48.1

Rank #57 Confirmed

Summary

  • They share 24 benchmarks with published results for both. Gemini 2.5 Flash-Lite scores higher in 0 categories and Kimi K2.5 in 10 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Kimi K2.5 leads 53.6 to 32.5.
  • The biggest single-benchmark swing is Fiction.LiveBench: 47.2% for Gemini 2.5 Flash-Lite and 86.1% for Kimi K2.5.
  • Gemini 2.5 Flash-Lite is cheaper at $0.10 / $0.40 per million input/output tokens, against $0.45 / $2.25 for Kimi K2.5.
  • Gemini 2.5 Flash-Lite accepts more context: 1.05M tokens versus 262K.
  • Kimi K2.5 has downloadable open weights; the other is API-only.

Side by side

Gemini 2.5 Flash-Lite and Kimi K2.5 specifications
Gemini 2.5 Flash-LiteKimi K2.5
ProviderGoogleMoonshot AI
Noometry Index37.048.1
Released2025-06-172026-01-27
WeightsProprietaryOpen
Context window1.05M262K
Max output66K262K
Input $ / M tokens$0.10$0.45
Output $ / M tokens$0.40$2.25
Results tracked3351

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2.5 leads

Gemini 2.5 Flash-Lite: 38.5 (#173), Kimi K2.5: 48.8 (#53)

Coding benchmarks
BenchmarkGemini 2.5 Flash-LiteKimi K2.5
WeirdML35.2%45.6%
LMArena Coding13731474
ALE-Bench325.9821.65
SWE-bench Verified—73.8%
SWE-bench Verified (bash only)—70.8%
LMArena WebDev—1437
SWE-bench Multilingual—67.3%
SciCode—49%

Agentic & Tool Use Kimi K2.5 leads

Gemini 2.5 Flash-Lite: 28.0 (#96), Kimi K2.5: 34.2 (#48)

Agentic & Tool Use benchmarks
BenchmarkGemini 2.5 Flash-LiteKimi K2.5
Terminal-Bench—43.2%
Berkeley Function Calling Leaderboard36.9%—
OSWorld—63.3%
Vending-Bench 2—1,198

Reasoning Kimi K2.5 leads

Gemini 2.5 Flash-Lite: 22.2 (#205), Kimi K2.5: 31.2 (#80)

Reasoning benchmarks
BenchmarkGemini 2.5 Flash-LiteKimi K2.5
Kagi LLM Benchmark40.5%78.5%
LMArena Hard Prompts13771453
Epoch Capabilities Index133.94148.03
ARC-AGI-2—11.8%
SimpleBench—46.8%
NYT Connections (extended)—69.9%
ARC-AGI-1—65.3%
CritPt—3.1%
Chess Puzzles—12%
EnigmaEval—3.4%
Thematic Generalization—69.4%
DTBench62.8%—
LMCA18.1%—

Math Kimi K2.5 leads

Gemini 2.5 Flash-Lite: 38.0 (#144), Kimi K2.5: 51.8 (#53)

Math benchmarks
BenchmarkGemini 2.5 Flash-LiteKimi K2.5
LMArena Math13731470
MathArena Final-Answer Competitions—62.3%
OTIS Mock AIME 2024-2025—92.2%
Omni-MATH48%—
FrontierMath (Feb 2025 set)—27.9%
FrontierMath Tier 4 (v1)—4.2%

Knowledge Kimi K2.5 leads

Gemini 2.5 Flash-Lite: 32.5 (#210), Kimi K2.5: 53.6 (#56)

Knowledge benchmarks
BenchmarkGemini 2.5 Flash-LiteKimi K2.5
Vectara Hallucination Rate3.3%14.2%
LMArena Expert13731466
GPQA Diamond—87.6%
Humanity's Last Exam—24.4%
SimpleQA Verified—34.3%
MMLU-Pro53.7%—
GPQA (HELM)30.9%—

Multimodal Kimi K2.5 leads

Gemini 2.5 Flash-Lite: 29.1 (#114), Kimi K2.5: 41.1 (#39)

Multimodal benchmarks
BenchmarkGemini 2.5 Flash-LiteKimi K2.5
LMArena Vision11981269
VPCT30%—
LMArena Document—1430

Multilingual Kimi K2.5 leads

Gemini 2.5 Flash-Lite: 49.3 (#134), Kimi K2.5: 53.9 (#53)

Multilingual benchmarks
BenchmarkGemini 2.5 Flash-LiteKimi K2.5
LMArena Non-English13691433
LMArena Chinese14041495
LMArena French13881454
LMArena German13891441
LMArena Japanese13591421
LMArena Korean13601410
LMArena Russian13731435
LMArena Spanish13961450

Instruction Following Kimi K2.5 leads

Gemini 2.5 Flash-Lite: 70.0 (#168), Kimi K2.5: 75.3 (#64)

Instruction Following benchmarks
BenchmarkGemini 2.5 Flash-LiteKimi K2.5
LMArena Instruction Following13671431
IFEval81%—

Long Context Kimi K2.5 leads

Gemini 2.5 Flash-Lite: 33.3 (#262), Kimi K2.5: 52.1 (#7)

Long Context benchmarks
BenchmarkGemini 2.5 Flash-LiteKimi K2.5
Fiction.LiveBench47.2%86.1%
LMArena Longer Query13731445
CL-bench—19.3%
CL-bench Life—13.2%

Writing & Preference Kimi K2.5 leads

Gemini 2.5 Flash-Lite: 56.8 (#135), Kimi K2.5: 65.1 (#53)

Writing & Preference benchmarks
BenchmarkGemini 2.5 Flash-LiteKimi K2.5
LMArena Text13791445
LMArena Creative Writing13671423
LMArena Multi-Turn13661444
EQ-Bench Creative Writing—1579
WildBench81.8%—

Frequently asked questions

Is Gemini 2.5 Flash-Lite better than Kimi K2.5?

Kimi K2.5 is the stronger model overall, scoring 48.1 to 37.0 on the Noometry Index. Gemini 2.5 Flash-Lite costs 5.1× less per token, which makes it the better buy when Kimi K2.5's lead doesn't matter for your workload.

Which is cheaper, Gemini 2.5 Flash-Lite or Kimi K2.5?

Gemini 2.5 Flash-Lite is cheaper. It lists at $0.10 per million input tokens and $0.40 per million output tokens; Kimi K2.5 lists at $0.45 and $2.25.

Is Gemini 2.5 Flash-Lite or Kimi K2.5 better for coding?

Kimi K2.5 scores higher on coding benchmarks: 48.8 versus 38.5 in the Noometry coding category.

Which has the bigger context window?

Gemini 2.5 Flash-Lite does, with 1.05M tokens against 262K.

How many benchmarks do Gemini 2.5 Flash-Lite and Kimi K2.5 share?

24 benchmarks have published results for both models. Gemini 2.5 Flash-Lite has 33 scored results on Noometry and Kimi K2.5 has 51.

Related comparisons

Go deeper