Model comparison

Kimi K2.5 vs Qwen2-72B

Kimi K2.5 is the stronger model overall, scoring 48.1 to 30.0 on the Noometry Index.

Last verified . 20 shared benchmarks.

Kimi K2.5 Moonshot AI

48.1

Rank #57 Confirmed

Qwen2-72B Alibaba (Qwen)

30.0

Rank #300 Confirmed

Summary

  • They share 20 benchmarks with published results for both. Kimi K2.5 scores higher in 9 categories and Qwen2-72B in 0 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Kimi K2.5 leads 53.6 to 21.2.
  • The biggest single-benchmark swing is GPQA Diamond: 87.6% for Kimi K2.5 and 40.8% for Qwen2-72B.

Side by side

Kimi K2.5 and Qwen2-72B specifications
Kimi K2.5Qwen2-72B
ProviderMoonshot AIAlibaba (Qwen)
Noometry Index48.130.0
Released2026-01-272024-06-07
WeightsOpenOpen
Context window262K—
Max output262K—
Input $ / M tokens$0.45—
Output $ / M tokens$2.25—
Results tracked5126

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2.5 leads

Kimi K2.5: 48.8 (#53), Qwen2-72B: 29.1 (#310)

Coding benchmarks
BenchmarkKimi K2.5Qwen2-72B
WeirdML45.6%11.3%
LMArena Coding14741196
SWE-bench Verified73.8%—
SWE-bench Verified (bash only)70.8%—
LMArena WebDev1437—
SWE-bench Multilingual67.3%—
SciCode49%—
BigCodeBench Instruct—38.5%
BigCodeBench Complete—54%
ALE-Bench821.65—

Agentic & Tool Use Kimi K2.5 leads

Kimi K2.5: 34.2 (#48), Qwen2-72B: 17.0 (#146)

Agentic & Tool Use benchmarks
BenchmarkKimi K2.5Qwen2-72B
Terminal-Bench43.2%—
TheAgentCompany—1.1%
OSWorld63.3%—
METR Time Horizons—29.9%
Vending-Bench 21,198—

Reasoning Kimi K2.5 leads

Kimi K2.5: 31.2 (#80), Qwen2-72B: 23.2 (#181)

Reasoning benchmarks
BenchmarkKimi K2.5Qwen2-72B
LMArena Hard Prompts14531191
Epoch Capabilities Index148.03125.28
ARC-AGI-211.8%—
SimpleBench46.8%—
Kagi LLM Benchmark78.5%—
NYT Connections (extended)69.9%—
ARC-AGI-165.3%—
CritPt3.1%—
Chess Puzzles12%—
EnigmaEval3.4%—
Thematic Generalization69.4%—

Math Kimi K2.5 leads

Kimi K2.5: 51.8 (#53), Qwen2-72B: 30.2 (#236)

Knowledge Kimi K2.5 leads

Kimi K2.5: 53.6 (#56), Qwen2-72B: 21.2 (#275)

Knowledge benchmarks
BenchmarkKimi K2.5Qwen2-72B
GPQA Diamond87.6%40.8%
LMArena Expert14661171
Humanity's Last Exam24.4%—
SimpleQA Verified34.3%—
Vectara Hallucination Rate14.2%—
MMLU—82.4%

Multimodal Not comparable

Kimi K2.5: 41.1 (#39), Qwen2-72B: —

Multimodal benchmarks
BenchmarkKimi K2.5Qwen2-72B
LMArena Vision1269—
LMArena Document1430—

Multilingual Kimi K2.5 leads

Kimi K2.5: 53.9 (#53), Qwen2-72B: 35.9 (#244)

Multilingual benchmarks
BenchmarkKimi K2.5Qwen2-72B
LMArena Non-English14331176
LMArena Chinese14951240
LMArena French14541170
LMArena German14411151
LMArena Japanese14211111
LMArena Korean14101083
LMArena Russian14351169
LMArena Spanish14501169

Instruction Following Kimi K2.5 leads

Kimi K2.5: 75.3 (#64), Qwen2-72B: 61.7 (#241)

Instruction Following benchmarks
BenchmarkKimi K2.5Qwen2-72B
LMArena Instruction Following14311181

Long Context Kimi K2.5 leads

Kimi K2.5: 52.1 (#7), Qwen2-72B: 36.1 (#235)

Long Context benchmarks
BenchmarkKimi K2.5Qwen2-72B
LMArena Longer Query14451192
Fiction.LiveBench86.1%—
CL-bench19.3%—
CL-bench Life13.2%—

Writing & Preference Kimi K2.5 leads

Kimi K2.5: 65.1 (#53), Qwen2-72B: 40.8 (#241)

Writing & Preference benchmarks
BenchmarkKimi K2.5Qwen2-72B
LMArena Text14451203
LMArena Creative Writing14231181
LMArena Multi-Turn14441196
EQ-Bench Creative Writing1579—

Frequently asked questions

Is Kimi K2.5 better than Qwen2-72B?

Kimi K2.5 is the stronger model overall, scoring 48.1 to 30.0 on the Noometry Index.

Is Kimi K2.5 or Qwen2-72B better for coding?

Kimi K2.5 scores higher on coding benchmarks: 48.8 versus 29.1 in the Noometry coding category.

How many benchmarks do Kimi K2.5 and Qwen2-72B share?

20 benchmarks have published results for both models. Kimi K2.5 has 51 scored results on Noometry and Qwen2-72B has 26.

Related comparisons

Go deeper