Model comparison

DeepSeek-R1 vs Kimi K2.7 Code

DeepSeek-R1 and Kimi K2.7 Code score almost the same on the Noometry Index (42.3 vs 43.3), so choose on price, context window or the category you care about most.

Last verified . 8 shared benchmarks.

DeepSeek-R1 DeepSeek

42.3

Rank #115 Confirmed

Kimi K2.7 Code Moonshot AI

43.3

Rank #94 Confirmed

Summary

  • They share 8 benchmarks with published results for both. DeepSeek-R1 scores higher in 2 categories and Kimi K2.7 Code in 3 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Kimi K2.7 Code leads 39.0 to 18.6.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 66.4% for DeepSeek-R1 and 95.6% for Kimi K2.7 Code.
  • DeepSeek-R1 is cheaper at $0.50 / $2.15 per million input/output tokens, against $0.95 / $4 for Kimi K2.7 Code.
  • Kimi K2.7 Code accepts more context: 262K tokens versus 164K.
  • Kimi K2.7 Code has downloadable open weights; the other is API-only.

Side by side

DeepSeek-R1 and Kimi K2.7 Code specifications
DeepSeek-R1Kimi K2.7 Code
ProviderDeepSeekMoonshot AI
Noometry Index42.343.3
Released2025-01-202026-06-12
WeightsProprietaryOpen
Context window164K262K
Max output64K262K
Input $ / M tokens$0.50$0.95
Output $ / M tokens$2.15$4
Results tracked5219

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-R1 leads

DeepSeek-R1: 46.3 (#68), Kimi K2.7 Code: 42.9 (#95)

Coding benchmarks
BenchmarkDeepSeek-R1Kimi K2.7 Code
SciCode35.7%47.5%
WeirdML41.6%54.1%
ALE-Bench804.12886.23
DeepSWE—30.5%
FrontierCode—30.1%
Aider Polyglot71.4%—
LMArena WebDev—1473
LiveBench Coding66.7%—
LMArena Coding1427—
AlgoTune1.7—

Agentic & Tool Use DeepSeek-R1 leads

DeepSeek-R1: 30.7 (#75), Kimi K2.7 Code: 24.0 (#122)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-R1Kimi K2.7 Code
APEX-Agents—37.6%
DeepResearch Bench35.1%—
BALROG34.9%—
GBAEval—0.9%
METR Time Horizons53.8%—
Vending-Bench 2—5,083

Reasoning Kimi K2.7 Code leads

DeepSeek-R1: 18.6 (#278), Kimi K2.7 Code: 39.0 (#61)

Reasoning benchmarks
BenchmarkDeepSeek-R1Kimi K2.7 Code
SimpleBench40.8%57.9%
CritPt1.1%10%
Epoch Capabilities Index141.29149.97
ARC-AGI-21.3%—
Kagi LLM Benchmark69.4%—
ARC-AGI-121.2%—
Chess Puzzles—21%
LiveBench Reasoning83.2%—
LMArena Hard Prompts1416—
LiveBench Data Analysis69.8%—
Surface Evolver Bench—48.8%
ForecastBench60—
LiveBench71.6%—

Math Kimi K2.7 Code leads

DeepSeek-R1: 43.8 (#79), Kimi K2.7 Code: 52.9 (#48)

Math benchmarks
BenchmarkDeepSeek-R1Kimi K2.7 Code
OTIS Mock AIME 2024-202566.4%95.6%
FrontierMath (Tiers 1-3)—54%
FrontierMath Tier 4—12.2%
Omni-MATH42.4%—
LiveBench Math80.7%—
LMArena Math1400—
MATH Level 596.6%—

Knowledge Kimi K2.7 Code leads

DeepSeek-R1: 44.5 (#87), Kimi K2.7 Code: 53.5 (#57)

Knowledge benchmarks
BenchmarkDeepSeek-R1Kimi K2.7 Code
GPQA Diamond76.3%87.9%
SimpleQA Verified—36.5%
MMLU-Pro79.3%—
Confabulations12.7%—
Vectara Hallucination Rate11.3%—
GPQA (HELM)66.6%—
LMArena Expert1394—

Multilingual Not comparable

DeepSeek-R1: 52.4 (#85), Kimi K2.7 Code: —

Multilingual benchmarks
BenchmarkDeepSeek-R1Kimi K2.7 Code
LMArena Non-English1412—
LMArena Chinese1442—
LMArena French1417—
LMArena German1404—
LMArena Japanese1391—
LMArena Korean1360—
LMArena Russian1423—
LMArena Spanish1411—

Instruction Following Not comparable

DeepSeek-R1: 72.0 (#143), Kimi K2.7 Code: —

Instruction Following benchmarks
BenchmarkDeepSeek-R1Kimi K2.7 Code
LiveBench Instruction Following80.5%—
IFEval78.4%—
LMArena Instruction Following1382—

Long Context Not comparable

DeepSeek-R1: 45.4 (#36), Kimi K2.7 Code: —

Long Context benchmarks
BenchmarkDeepSeek-R1Kimi K2.7 Code
Fiction.LiveBench75%—
LMArena Longer Query1391—

Writing & Preference Not comparable

DeepSeek-R1: 61.4 (#88), Kimi K2.7 Code: —

Writing & Preference benchmarks
BenchmarkDeepSeek-R1Kimi K2.7 Code
LMArena Text1428—
LMArena Creative Writing1405—
Short-Story Creative Writing83%—
EQ-Bench Creative Writing1500—
WildBench82.8%—
LMArena Multi-Turn1405—
LiveBench Language48.5%—

Frequently asked questions

Is DeepSeek-R1 better than Kimi K2.7 Code?

DeepSeek-R1 and Kimi K2.7 Code score almost the same on the Noometry Index (42.3 vs 43.3), so choose on price, context window or the category you care about most.

Which is cheaper, DeepSeek-R1 or Kimi K2.7 Code?

DeepSeek-R1 is cheaper. It lists at $0.50 per million input tokens and $2.15 per million output tokens; Kimi K2.7 Code lists at $0.95 and $4.

Is DeepSeek-R1 or Kimi K2.7 Code better for coding?

DeepSeek-R1 scores higher on coding benchmarks: 46.3 versus 42.9 in the Noometry coding category.

Which has the bigger context window?

Kimi K2.7 Code does, with 262K tokens against 164K.

How many benchmarks do DeepSeek-R1 and Kimi K2.7 Code share?

8 benchmarks have published results for both models. DeepSeek-R1 has 52 scored results on Noometry and Kimi K2.7 Code has 19.

Related comparisons

Go deeper