Model comparison

DeepSeek-V3.1 vs Kimi K2.5 Instant

DeepSeek-V3.1 and Kimi K2.5 Instant score almost the same on the Noometry Index (42.8 vs 43.6), so choose on price, context window or the category you care about most.

Last verified . 16 shared benchmarks.

DeepSeek-V3.1 DeepSeek

42.8

Rank #108 Confirmed

Kimi K2.5 Instant Moonshot AI

43.6

Rank #89 Confirmed

Summary

  • They share 16 benchmarks with published results for both. DeepSeek-V3.1 scores higher in 1 category and Kimi K2.5 Instant in 7 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Kimi K2.5 Instant leads 43.9 to 36.3.

Side by side

DeepSeek-V3.1 and Kimi K2.5 Instant specifications
DeepSeek-V3.1Kimi K2.5 Instant
ProviderDeepSeekMoonshot AI
Noometry Index42.843.6
Released2025-08-21—
WeightsOpenOpen
Context window164K—
Max output8K—
Input $ / M tokens$0.25—
Output $ / M tokens$0.95—
Results tracked2718

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2.5 Instant leads

DeepSeek-V3.1: 40.3 (#144), Kimi K2.5 Instant: 42.6 (#97)

Coding benchmarks
BenchmarkDeepSeek-V3.1Kimi K2.5 Instant
LMArena Coding14171484
LMArena WebDev—1404
WeirdML38.4%—

Reasoning Kimi K2.5 Instant leads

DeepSeek-V3.1: 27.9 (#110), Kimi K2.5 Instant: 29.7 (#90)

Reasoning benchmarks
BenchmarkDeepSeek-V3.1Kimi K2.5 Instant
LMArena Hard Prompts14171443
SimpleBench40%—
Kagi LLM Benchmark53.2%—
DTBench82.7%—
LMCA24.3%—
Epoch Capabilities Index139.92—
ForecastBench58—

Math Too close to call

DeepSeek-V3.1: 38.9 (#122), Kimi K2.5 Instant: 39.4 (#105)

Math benchmarks
BenchmarkDeepSeek-V3.1Kimi K2.5 Instant
LMArena Math14201442

Knowledge DeepSeek-V3.1 leads

DeepSeek-V3.1: 43.7 (#90), Kimi K2.5 Instant: 40.2 (#123)

Knowledge benchmarks
BenchmarkDeepSeek-V3.1Kimi K2.5 Instant
LMArena Expert14051440
Vectara Hallucination Rate5.5%—

Multimodal Not comparable

DeepSeek-V3.1: —, Kimi K2.5 Instant: 40.2 (#50)

Multimodal benchmarks
BenchmarkDeepSeek-V3.1Kimi K2.5 Instant
LMArena Vision—1254

Multilingual Too close to call

DeepSeek-V3.1: 51.6 (#106), Kimi K2.5 Instant: 52.0 (#94)

Multilingual benchmarks
BenchmarkDeepSeek-V3.1Kimi K2.5 Instant
LMArena Non-English14001406
LMArena Chinese14691449
LMArena French14471403
LMArena German14111413
LMArena Korean13371378
LMArena Russian14051404
LMArena Spanish14311447
LMArena Japanese1378—

Instruction Following Kimi K2.5 Instant leads

DeepSeek-V3.1: 73.9 (#110), Kimi K2.5 Instant: 75.3 (#65)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.1Kimi K2.5 Instant
LMArena Instruction Following14001430

Long Context Kimi K2.5 Instant leads

DeepSeek-V3.1: 36.3 (#232), Kimi K2.5 Instant: 43.9 (#83)

Long Context benchmarks
BenchmarkDeepSeek-V3.1Kimi K2.5 Instant
LMArena Longer Query14221435
Fiction.LiveBench52.8%—

Writing & Preference Too close to call

DeepSeek-V3.1: 60.3 (#98), Kimi K2.5 Instant: 60.6 (#95)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.1Kimi K2.5 Instant
LMArena Text14201420
LMArena Creative Writing14011381
LMArena Multi-Turn14081427
EQ-Bench Creative Writing1436—

Frequently asked questions

Is DeepSeek-V3.1 better than Kimi K2.5 Instant?

DeepSeek-V3.1 and Kimi K2.5 Instant score almost the same on the Noometry Index (42.8 vs 43.6), so choose on price, context window or the category you care about most.

Is DeepSeek-V3.1 or Kimi K2.5 Instant better for coding?

Kimi K2.5 Instant scores higher on coding benchmarks: 42.6 versus 40.3 in the Noometry coding category.

How many benchmarks do DeepSeek-V3.1 and Kimi K2.5 Instant share?

16 benchmarks have published results for both models. DeepSeek-V3.1 has 27 scored results on Noometry and Kimi K2.5 Instant has 18.

Related comparisons

Go deeper