Model comparison

DeepSeek V4 Flash vs Kimi K2.5

DeepSeek V4 Flash is the stronger model overall, scoring 53.6 to 48.1 on the Noometry Index.

Last verified . 34 shared benchmarks.

DeepSeek V4 Flash DeepSeek

53.6

Rank #35 Confirmed

Kimi K2.5 Moonshot AI

48.1

Rank #57 Confirmed

Summary

  • They share 34 benchmarks with published results for both. DeepSeek V4 Flash scores higher in 3 categories and Kimi K2.5 in 5 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where DeepSeek V4 Flash leads 53.7 to 31.2.
  • The biggest single-benchmark swing is ARC-AGI-2: 61.4% for DeepSeek V4 Flash and 11.8% for Kimi K2.5.
  • DeepSeek V4 Flash is cheaper at $0.15 / $0.60 per million input/output tokens, against $0.45 / $2.25 for Kimi K2.5.
  • DeepSeek V4 Flash accepts more context: 1M tokens versus 262K.

Side by side

DeepSeek V4 Flash and Kimi K2.5 specifications
DeepSeek V4 FlashKimi K2.5
ProviderDeepSeekMoonshot AI
Noometry Index53.648.1
Released2026-04-242026-01-27
WeightsOpenOpen
Context window1M262K
Max output393K262K
Input $ / M tokens$0.15$0.45
Output $ / M tokens$0.60$2.25
Results tracked4151

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

DeepSeek V4 Flash: 47.9 (#59), Kimi K2.5: 48.8 (#53)

Coding benchmarks
BenchmarkDeepSeek V4 FlashKimi K2.5
LMArena WebDev15821437
SciCode49.9%49%
WeirdML63%45.6%
LMArena Coding14571474
ALE-Bench1,306821.65
SWE-bench Verified—73.8%
FrontierCode18.8%—
SWE-bench Verified (bash only)—70.8%
SWE-bench Multilingual—67.3%

Agentic & Tool Use Not comparable

DeepSeek V4 Flash: —, Kimi K2.5: 34.2 (#48)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek V4 FlashKimi K2.5
Terminal-Bench—43.2%
OSWorld—63.3%
Vending-Bench 2—1,198

Reasoning DeepSeek V4 Flash leads

DeepSeek V4 Flash: 53.7 (#30), Kimi K2.5: 31.2 (#80)

Reasoning benchmarks
BenchmarkDeepSeek V4 FlashKimi K2.5
ARC-AGI-261.4%11.8%
SimpleBench61.1%46.8%
Kagi LLM Benchmark52.2%78.5%
NYT Connections (extended)89.6%69.9%
ARC-AGI-189%65.3%
CritPt16.6%3.1%
Chess Puzzles33%12%
LMArena Hard Prompts14441453
Epoch Capabilities Index154.49148.03
EnigmaEval—3.4%
Thematic Generalization—69.4%
Mystery Game Puzzles34%—
DTBench90.9%—
LMCA41.7%—

Math DeepSeek V4 Flash leads

DeepSeek V4 Flash: 60.3 (#37), Kimi K2.5: 51.8 (#53)

Knowledge DeepSeek V4 Flash leads

DeepSeek V4 Flash: 55.4 (#48), Kimi K2.5: 53.6 (#56)

Knowledge benchmarks
BenchmarkDeepSeek V4 FlashKimi K2.5
GPQA Diamond91%87.6%
SimpleQA Verified33.6%34.3%
LMArena Expert14411466
Humanity's Last Exam—24.4%
Vectara Hallucination Rate—14.2%

Multimodal Not comparable

DeepSeek V4 Flash: —, Kimi K2.5: 41.1 (#39)

Multimodal benchmarks
BenchmarkDeepSeek V4 FlashKimi K2.5
LMArena Vision—1269
LMArena Document—1430

Multilingual Too close to call

DeepSeek V4 Flash: 53.0 (#72), Kimi K2.5: 53.9 (#53)

Multilingual benchmarks
BenchmarkDeepSeek V4 FlashKimi K2.5
LMArena Non-English14201433
LMArena Chinese14681495
LMArena French14391454
LMArena German14181441
LMArena Japanese14061421
LMArena Korean13841410
LMArena Russian14281435
LMArena Spanish14361450

Instruction Following Too close to call

DeepSeek V4 Flash: 74.9 (#81), Kimi K2.5: 75.3 (#64)

Instruction Following benchmarks
BenchmarkDeepSeek V4 FlashKimi K2.5
LMArena Instruction Following14211431

Long Context Kimi K2.5 leads

DeepSeek V4 Flash: 43.8 (#85), Kimi K2.5: 52.1 (#7)

Long Context benchmarks
BenchmarkDeepSeek V4 FlashKimi K2.5
LMArena Longer Query14341445
Fiction.LiveBench—86.1%
CL-bench—19.3%
CL-bench Life—13.2%

Writing & Preference Kimi K2.5 leads

DeepSeek V4 Flash: 63.8 (#61), Kimi K2.5: 65.1 (#53)

Writing & Preference benchmarks
BenchmarkDeepSeek V4 FlashKimi K2.5
LMArena Text14321445
LMArena Creative Writing14031423
EQ-Bench Creative Writing15591579
LMArena Multi-Turn14491444

Frequently asked questions

Is DeepSeek V4 Flash better than Kimi K2.5?

DeepSeek V4 Flash is the stronger model overall, scoring 53.6 to 48.1 on the Noometry Index.

Which is cheaper, DeepSeek V4 Flash or Kimi K2.5?

DeepSeek V4 Flash is cheaper. It lists at $0.15 per million input tokens and $0.60 per million output tokens; Kimi K2.5 lists at $0.45 and $2.25.

Is DeepSeek V4 Flash or Kimi K2.5 better for coding?

They score almost the same on coding (47.9 vs 48.8); test both on your own repository before choosing.

Which has the bigger context window?

DeepSeek V4 Flash does, with 1M tokens against 262K.

How many benchmarks do DeepSeek V4 Flash and Kimi K2.5 share?

34 benchmarks have published results for both models. DeepSeek V4 Flash has 41 scored results on Noometry and Kimi K2.5 has 51.

Related comparisons

Go deeper