Model comparison

DeepSeek-V3 vs Kimi K2 (Jul 2025)

Kimi K2 (Jul 2025) is the stronger model overall, scoring 41.2 to 39.5 on the Noometry Index. DeepSeek-V3 costs 2.5× less per token, which makes it the better buy when Kimi K2 (Jul 2025)'s lead doesn't matter for your workload.

Last verified . 35 shared benchmarks.

DeepSeek-V3 DeepSeek

39.5

Rank #166 Confirmed

Kimi K2 (Jul 2025) Moonshot AI

41.2

Rank #140 Confirmed

Summary

  • They share 35 benchmarks with published results for both. DeepSeek-V3 scores higher in 2 categories and Kimi K2 (Jul 2025) in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where Kimi K2 (Jul 2025) leads 42.7 to 32.1.
  • The biggest single-benchmark swing is Omni-MATH: 40.3% for DeepSeek-V3 and 65.4% for Kimi K2 (Jul 2025).
  • DeepSeek-V3 is cheaper at $0.24 / $0.90 per million input/output tokens, against $0.57 / $2.30 for Kimi K2 (Jul 2025).
  • Kimi K2 (Jul 2025) accepts more context: 262K tokens versus 164K.

Side by side

DeepSeek-V3 and Kimi K2 (Jul 2025) specifications
DeepSeek-V3Kimi K2 (Jul 2025)
ProviderDeepSeekMoonshot AI
Noometry Index39.541.2
Released2024-12-262025-07-12
WeightsOpenOpen
Context window164K262K
Max output164K262K
Input $ / M tokens$0.24$0.57
Output $ / M tokens$0.90$2.30
Results tracked6042

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

DeepSeek-V3: 42.3 (#106), Kimi K2 (Jul 2025): 42.4 (#102)

Coding benchmarks
BenchmarkDeepSeek-V3Kimi K2 (Jul 2025)
Aider Polyglot55.1%59.1%
WeirdML36.1%42.8%
LMArena Coding13681399
SWE-bench Verified (bash only)—63.4%
SciCode35.8%—
GSO—4.9%
BigCodeBench Instruct50%—
LiveBench Coding70.9%—
BigCodeBench Complete62.2%—
ALE-Bench—597.5
HumanEval+86.6%—
MBPP+73%—

Agentic & Tool Use Not comparable

DeepSeek-V3: —, Kimi K2 (Jul 2025): 32.4 (#64)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3Kimi K2 (Jul 2025)
METR Time Horizons49.6%59.2%
Terminal-Bench—35.7%
Berkeley Function Calling Leaderboard—59.1%

Reasoning Kimi K2 (Jul 2025) leads

DeepSeek-V3: 20.5 (#236), Kimi K2 (Jul 2025): 23.3 (#179)

Reasoning benchmarks
BenchmarkDeepSeek-V3Kimi K2 (Jul 2025)
SimpleBench27.2%26.3%
Kagi LLM Benchmark52.3%64.4%
LMArena Hard Prompts13651384
Epoch Capabilities Index135.94146.01
ForecastBench59.160.2
CritPt0%—
LiveBench Reasoning65.8%—
DTBench64.8%—
LiveBench Data Analysis60.9%—
LMCA15.5%—
BIG-Bench Hard87.5%—
HellaSwag88.9%—
LiveBench66.9%—
PIQA84.7%—
WinoGrande85.2%—

Math Kimi K2 (Jul 2025) leads

DeepSeek-V3: 32.1 (#219), Kimi K2 (Jul 2025): 42.7 (#83)

Math benchmarks
BenchmarkDeepSeek-V3Kimi K2 (Jul 2025)
Omni-MATH40.3%65.4%
LMArena Math13731397
FrontierMath (Feb 2025 set)1.7%21.4%
OTIS Mock AIME 2024-202537.8%—
LiveBench Math73.5%—
MATH Level 575.5%—
FrontierMath Tier 4 (v1)—0%

Knowledge Too close to call

DeepSeek-V3: 37.5 (#155), Kimi K2 (Jul 2025): 37.3 (#157)

Knowledge benchmarks
BenchmarkDeepSeek-V3Kimi K2 (Jul 2025)
MMLU-Pro72.3%81.9%
Confabulations26.1%20.4%
Vectara Hallucination Rate6.1%17.9%
GPQA (HELM)53.8%65.3%
LMArena Expert13511365
GPQA Diamond67.6%—
ARC (AI2) Challenge95.3%—
MMLU87.2%—
TriviaQA82.9%—

Multilingual Kimi K2 (Jul 2025) leads

DeepSeek-V3: 48.5 (#143), Kimi K2 (Jul 2025): 49.6 (#130)

Multilingual benchmarks
BenchmarkDeepSeek-V3Kimi K2 (Jul 2025)
LMArena Non-English13581372
LMArena Chinese13911415
LMArena French13851379
LMArena German13741387
LMArena Japanese13331349
LMArena Korean13191325
LMArena Russian13731385
LMArena Spanish13581386

Instruction Following DeepSeek-V3 leads

DeepSeek-V3: 72.8 (#130), Kimi K2 (Jul 2025): 71.1 (#156)

Instruction Following benchmarks
BenchmarkDeepSeek-V3Kimi K2 (Jul 2025)
IFEval83.2%85%
LMArena Instruction Following13451348
LiveBench Instruction Following81.5%—

Long Context Kimi K2 (Jul 2025) leads

DeepSeek-V3: 34.0 (#253), Kimi K2 (Jul 2025): 41.2 (#145)

Long Context benchmarks
BenchmarkDeepSeek-V3Kimi K2 (Jul 2025)
Fiction.LiveBench50%66.7%
LMArena Longer Query13521353
CL-bench—17.6%

Writing & Preference Kimi K2 (Jul 2025) leads

DeepSeek-V3: 57.4 (#130), Kimi K2 (Jul 2025): 62.3 (#78)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3Kimi K2 (Jul 2025)
LMArena Text13751380
LMArena Creative Writing13641350
Short-Story Creative Writing77%85.6%
EQ-Bench Creative Writing14721666
WildBench83%86.2%
LMArena Multi-Turn13891371
LiveBench Language49.1%—

Frequently asked questions

Is DeepSeek-V3 better than Kimi K2 (Jul 2025)?

Kimi K2 (Jul 2025) is the stronger model overall, scoring 41.2 to 39.5 on the Noometry Index. DeepSeek-V3 costs 2.5× less per token, which makes it the better buy when Kimi K2 (Jul 2025)'s lead doesn't matter for your workload.

Which is cheaper, DeepSeek-V3 or Kimi K2 (Jul 2025)?

DeepSeek-V3 is cheaper. It lists at $0.24 per million input tokens and $0.90 per million output tokens; Kimi K2 (Jul 2025) lists at $0.57 and $2.30.

Is DeepSeek-V3 or Kimi K2 (Jul 2025) better for coding?

They score almost the same on coding (42.3 vs 42.4); test both on your own repository before choosing.

Which has the bigger context window?

Kimi K2 (Jul 2025) does, with 262K tokens against 164K.

How many benchmarks do DeepSeek-V3 and Kimi K2 (Jul 2025) share?

35 benchmarks have published results for both models. DeepSeek-V3 has 60 scored results on Noometry and Kimi K2 (Jul 2025) has 42.

Related comparisons

Go deeper