Model comparison

DeepSeek-R1-Distill-Llama-70B vs Kimi K2.7 Code

Kimi K2.7 Code is the stronger model overall, scoring 43.3 to 37.8 on the Noometry Index.

Last verified . 2 shared benchmarks.

Kimi K2.7 Code Moonshot AI

43.3

Rank #94 Confirmed

Summary

  • They share 2 benchmarks with published results for both. DeepSeek-R1-Distill-Llama-70B scores higher in 0 categories and Kimi K2.7 Code in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Kimi K2.7 Code leads 53.5 to 30.7.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 51.4% for DeepSeek-R1-Distill-Llama-70B and 95.6% for Kimi K2.7 Code.

Side by side

DeepSeek-R1-Distill-Llama-70B and Kimi K2.7 Code specifications
DeepSeek-R1-Distill-Llama-70BKimi K2.7 Code
ProviderDeepSeekMoonshot AI
Noometry Index37.843.3
Released2025-01-202026-06-12
WeightsOpenOpen
Context window—262K
Max output—262K
Input $ / M tokens—$0.95
Output $ / M tokens—$4
Results tracked1319

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2.7 Code leads

DeepSeek-R1-Distill-Llama-70B: 36.8 (#202), Kimi K2.7 Code: 42.9 (#95)

Coding benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BKimi K2.7 Code
DeepSWE—30.5%
FrontierCode—30.1%
LMArena WebDev—1473
SciCode—47.5%
WeirdML—54.1%
BigCodeBench Instruct35.3%—
LiveBench Coding51.6%—
BigCodeBench Complete49.9%—
ALE-Bench—886.23

Agentic & Tool Use Not comparable

DeepSeek-R1-Distill-Llama-70B: —, Kimi K2.7 Code: 24.0 (#122)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BKimi K2.7 Code
APEX-Agents—37.6%
GBAEval—0.9%
Vending-Bench 2—5,083

Reasoning Kimi K2.7 Code leads

DeepSeek-R1-Distill-Llama-70B: 24.9 (#156), Kimi K2.7 Code: 39.0 (#61)

Reasoning benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BKimi K2.7 Code
SimpleBench—57.9%
Kagi LLM Benchmark52.3%—
CritPt—10%
Chess Puzzles—21%
LiveBench Reasoning67.6%—
LiveBench Data Analysis55.9%—
Surface Evolver Bench—48.8%
Epoch Capabilities Index—149.97
LiveBench54.5%—

Math Kimi K2.7 Code leads

DeepSeek-R1-Distill-Llama-70B: 36.0 (#176), Kimi K2.7 Code: 52.9 (#48)

Math benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BKimi K2.7 Code
OTIS Mock AIME 2024-202551.4%95.6%
FrontierMath (Tiers 1-3)—54%
FrontierMath Tier 4—12.2%
LiveBench Math58.1%—
MATH Level 589.9%—

Knowledge Kimi K2.7 Code leads

DeepSeek-R1-Distill-Llama-70B: 30.7 (#225), Kimi K2.7 Code: 53.5 (#57)

Knowledge benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BKimi K2.7 Code
GPQA Diamond55.7%87.9%
SimpleQA Verified—36.5%

Instruction Following Not comparable

DeepSeek-R1-Distill-Llama-70B: 68.2 (#190), Kimi K2.7 Code: —

Instruction Following benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BKimi K2.7 Code
LiveBench Instruction Following69.9%—

Writing & Preference Not comparable

DeepSeek-R1-Distill-Llama-70B: 49.0 (#194), Kimi K2.7 Code: —

Writing & Preference benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BKimi K2.7 Code
LiveBench Language23.8%—

Frequently asked questions

Is DeepSeek-R1-Distill-Llama-70B better than Kimi K2.7 Code?

Kimi K2.7 Code is the stronger model overall, scoring 43.3 to 37.8 on the Noometry Index.

Is DeepSeek-R1-Distill-Llama-70B or Kimi K2.7 Code better for coding?

Kimi K2.7 Code scores higher on coding benchmarks: 42.9 versus 36.8 in the Noometry coding category.

How many benchmarks do DeepSeek-R1-Distill-Llama-70B and Kimi K2.7 Code share?

2 benchmarks have published results for both models. DeepSeek-R1-Distill-Llama-70B has 13 scored results on Noometry and Kimi K2.7 Code has 19.

Related comparisons

Go deeper