Model comparison

Kimi K2.5 Instant vs Qwen2.5 Plus 1127

Kimi K2.5 Instant is the stronger model overall, scoring 43.6 to 38.8 on the Noometry Index.

Last verified . 13 shared benchmarks.

Kimi K2.5 Instant Moonshot AI

43.6

Rank #89 Confirmed

Qwen2.5 Plus 1127 Alibaba (Qwen)

38.8

Rank #181 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Kimi K2.5 Instant scores higher in 8 categories and Qwen2.5 Plus 1127 in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Kimi K2.5 Instant leads 60.6 to 49.4.
  • Kimi K2.5 Instant has downloadable open weights; the other is API-only.

Side by side

Kimi K2.5 Instant and Qwen2.5 Plus 1127 specifications
Kimi K2.5 InstantQwen2.5 Plus 1127
ProviderMoonshot AIAlibaba (Qwen)
Noometry Index43.638.8
Released——
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1814

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2.5 Instant leads

Kimi K2.5 Instant: 42.6 (#97), Qwen2.5 Plus 1127: 38.5 (#175)

Coding benchmarks
BenchmarkKimi K2.5 InstantQwen2.5 Plus 1127
LMArena Coding14841314
LMArena WebDev1404—

Reasoning Kimi K2.5 Instant leads

Kimi K2.5 Instant: 29.7 (#90), Qwen2.5 Plus 1127: 25.9 (#141)

Reasoning benchmarks
BenchmarkKimi K2.5 InstantQwen2.5 Plus 1127
LMArena Hard Prompts14431299

Math Kimi K2.5 Instant leads

Kimi K2.5 Instant: 39.4 (#105), Qwen2.5 Plus 1127: 36.1 (#174)

Math benchmarks
BenchmarkKimi K2.5 InstantQwen2.5 Plus 1127
LMArena Math14421298

Knowledge Kimi K2.5 Instant leads

Kimi K2.5 Instant: 40.2 (#123), Qwen2.5 Plus 1127: 35.5 (#183)

Knowledge benchmarks
BenchmarkKimi K2.5 InstantQwen2.5 Plus 1127
LMArena Expert14401289

Multimodal Not comparable

Kimi K2.5 Instant: 40.2 (#50), Qwen2.5 Plus 1127: —

Multimodal benchmarks
BenchmarkKimi K2.5 InstantQwen2.5 Plus 1127
LMArena Vision1254—

Multilingual Kimi K2.5 Instant leads

Kimi K2.5 Instant: 52.0 (#94), Qwen2.5 Plus 1127: 41.9 (#201)

Multilingual benchmarks
BenchmarkKimi K2.5 InstantQwen2.5 Plus 1127
LMArena Non-English14061265
LMArena Chinese14491314
LMArena German14131231
LMArena Russian14041271
LMArena French1403—
LMArena Japanese—1207
LMArena Korean1378—
LMArena Spanish1447—

Instruction Following Kimi K2.5 Instant leads

Kimi K2.5 Instant: 75.3 (#65), Qwen2.5 Plus 1127: 67.2 (#199)

Instruction Following benchmarks
BenchmarkKimi K2.5 InstantQwen2.5 Plus 1127
LMArena Instruction Following14301275

Long Context Kimi K2.5 Instant leads

Kimi K2.5 Instant: 43.9 (#83), Qwen2.5 Plus 1127: 39.2 (#184)

Long Context benchmarks
BenchmarkKimi K2.5 InstantQwen2.5 Plus 1127
LMArena Longer Query14351292

Writing & Preference Kimi K2.5 Instant leads

Kimi K2.5 Instant: 60.6 (#95), Qwen2.5 Plus 1127: 49.4 (#192)

Writing & Preference benchmarks
BenchmarkKimi K2.5 InstantQwen2.5 Plus 1127
LMArena Text14201299
LMArena Creative Writing13811262
LMArena Multi-Turn14271299

Frequently asked questions

Is Kimi K2.5 Instant better than Qwen2.5 Plus 1127?

Kimi K2.5 Instant is the stronger model overall, scoring 43.6 to 38.8 on the Noometry Index.

Is Kimi K2.5 Instant or Qwen2.5 Plus 1127 better for coding?

Kimi K2.5 Instant scores higher on coding benchmarks: 42.6 versus 38.5 in the Noometry coding category.

How many benchmarks do Kimi K2.5 Instant and Qwen2.5 Plus 1127 share?

13 benchmarks have published results for both models. Kimi K2.5 Instant has 18 scored results on Noometry and Qwen2.5 Plus 1127 has 14.

Related comparisons

Go deeper