Model comparison

DeepSeek V4 Flash vs Kimi K2 Thinking Turbo

DeepSeek V4 Flash is the stronger model overall, scoring 53.6 to 45.8 on the Noometry Index.

Last verified . 21 shared benchmarks.

DeepSeek V4 Flash DeepSeek

53.6

Rank #35 Confirmed

Kimi K2 Thinking Turbo Moonshot AI

45.8

Rank #70 Confirmed

Summary

  • They share 21 benchmarks with published results for both. DeepSeek V4 Flash scores higher in 8 categories and Kimi K2 Thinking Turbo in 0 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where DeepSeek V4 Flash leads 53.7 to 32.5.
  • The biggest single-benchmark swing is Chess Puzzles: 33% for DeepSeek V4 Flash and 20% for Kimi K2 Thinking Turbo.

Side by side

DeepSeek V4 Flash and Kimi K2 Thinking Turbo specifications
DeepSeek V4 FlashKimi K2 Thinking Turbo
ProviderDeepSeekMoonshot AI
Noometry Index53.645.8
Released2026-04-242025-11-06
WeightsOpenOpen
Context window1M—
Max output393K—
Input $ / M tokens$0.15—
Output $ / M tokens$0.60—
Results tracked4121

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek V4 Flash leads

DeepSeek V4 Flash: 47.9 (#59), Kimi K2 Thinking Turbo: 38.4 (#178)

Coding benchmarks
BenchmarkDeepSeek V4 FlashKimi K2 Thinking Turbo
LMArena WebDev15821322
LMArena Coding14571454
FrontierCode18.8%—
SciCode49.9%—
WeirdML63%—
ALE-Bench1,306—

Reasoning DeepSeek V4 Flash leads

DeepSeek V4 Flash: 53.7 (#30), Kimi K2 Thinking Turbo: 32.5 (#75)

Reasoning benchmarks
BenchmarkDeepSeek V4 FlashKimi K2 Thinking Turbo
Chess Puzzles33%20%
LMArena Hard Prompts14441428
ARC-AGI-261.4%—
SimpleBench61.1%—
Kagi LLM Benchmark52.2%—
NYT Connections (extended)89.6%—
ARC-AGI-189%—
CritPt16.6%—
Mystery Game Puzzles34%—
DTBench90.9%—
LMCA41.7%—
Epoch Capabilities Index154.49—

Math DeepSeek V4 Flash leads

DeepSeek V4 Flash: 60.3 (#37), Kimi K2 Thinking Turbo: 47.4 (#68)

Math benchmarks
BenchmarkDeepSeek V4 FlashKimi K2 Thinking Turbo
OTIS Mock AIME 2024-202594.4%83.1%
LMArena Math14271429
FrontierMath (Tiers 1-3)57.5%—
FrontierMath Tier 424.4%—
MathArena Final-Answer Competitions76.5%—
ProofBench56%—

Knowledge DeepSeek V4 Flash leads

DeepSeek V4 Flash: 55.4 (#48), Kimi K2 Thinking Turbo: 50.9 (#69)

Knowledge benchmarks
BenchmarkDeepSeek V4 FlashKimi K2 Thinking Turbo
GPQA Diamond91%84.2%
LMArena Expert14411439
SimpleQA Verified33.6%—

Multilingual DeepSeek V4 Flash leads

DeepSeek V4 Flash: 53.0 (#72), Kimi K2 Thinking Turbo: 51.4 (#109)

Multilingual benchmarks
BenchmarkDeepSeek V4 FlashKimi K2 Thinking Turbo
LMArena Non-English14201398
LMArena Chinese14681456
LMArena French14391425
LMArena German14181390
LMArena Japanese14061357
LMArena Korean13841330
LMArena Russian14281391
LMArena Spanish14361406

Instruction Following Too close to call

DeepSeek V4 Flash: 74.9 (#81), Kimi K2 Thinking Turbo: 74.0 (#109)

Instruction Following benchmarks
BenchmarkDeepSeek V4 FlashKimi K2 Thinking Turbo
LMArena Instruction Following14211403

Long Context Too close to call

DeepSeek V4 Flash: 43.8 (#85), Kimi K2 Thinking Turbo: 43.2 (#102)

Long Context benchmarks
BenchmarkDeepSeek V4 FlashKimi K2 Thinking Turbo
LMArena Longer Query14341415

Writing & Preference DeepSeek V4 Flash leads

DeepSeek V4 Flash: 63.8 (#61), Kimi K2 Thinking Turbo: 60.0 (#104)

Writing & Preference benchmarks
BenchmarkDeepSeek V4 FlashKimi K2 Thinking Turbo
LMArena Text14321415
LMArena Creative Writing14031374
LMArena Multi-Turn14491414
EQ-Bench Creative Writing1559—

Frequently asked questions

Is DeepSeek V4 Flash better than Kimi K2 Thinking Turbo?

DeepSeek V4 Flash is the stronger model overall, scoring 53.6 to 45.8 on the Noometry Index.

Is DeepSeek V4 Flash or Kimi K2 Thinking Turbo better for coding?

DeepSeek V4 Flash scores higher on coding benchmarks: 47.9 versus 38.4 in the Noometry coding category.

How many benchmarks do DeepSeek V4 Flash and Kimi K2 Thinking Turbo share?

21 benchmarks have published results for both models. DeepSeek V4 Flash has 41 scored results on Noometry and Kimi K2 Thinking Turbo has 21.

Related comparisons

Go deeper