Model comparison

Kimi K2.5 Instant vs Qwen3 14B

Kimi K2.5 Instant is the stronger model overall, scoring 43.6 to 35.5 on the Noometry Index.

Last verified . 0 shared benchmarks.

Kimi K2.5 Instant Moonshot AI

43.6

Rank #89 Confirmed

Qwen3 14B Alibaba (Qwen)

35.5

Rank #225 Confirmed

Summary

  • The widest gap is in reasoning, where Kimi K2.5 Instant leads 29.7 to 18.5.

Side by side

Kimi K2.5 Instant and Qwen3 14B specifications
Kimi K2.5 InstantQwen3 14B
ProviderMoonshot AIAlibaba (Qwen)
Noometry Index43.635.5
Released—2025-04
WeightsOpenOpen
Context window—131K
Max output—8K
Input $ / M tokens—$0.35
Output $ / M tokens—$1.40
Results tracked1812

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2.5 Instant leads

Kimi K2.5 Instant: 42.6 (#97), Qwen3 14B: 37.3 (#195)

Coding benchmarks
BenchmarkKimi K2.5 InstantQwen3 14B
LMArena WebDev1404—
SciCode—31.6%
LMArena Coding1484—

Agentic & Tool Use Not comparable

Kimi K2.5 Instant: —, Qwen3 14B: 29.6 (#83)

Agentic & Tool Use benchmarks
BenchmarkKimi K2.5 InstantQwen3 14B
Berkeley Function Calling Leaderboard—41%

Reasoning Kimi K2.5 Instant leads

Kimi K2.5 Instant: 29.7 (#90), Qwen3 14B: 18.5 (#280)

Reasoning benchmarks
BenchmarkKimi K2.5 InstantQwen3 14B
Kagi LLM Benchmark—49.1%
CritPt—0%
Chess Puzzles—4%
LMArena Hard Prompts1443—
DTBench—64%
LMCA—18.2%
Epoch Capabilities Index—138.23

Math Too close to call

Kimi K2.5 Instant: 39.4 (#105), Qwen3 14B: 38.6 (#133)

Math benchmarks
BenchmarkKimi K2.5 InstantQwen3 14B
OTIS Mock AIME 2024-2025—66.4%
LMArena Math1442—

Knowledge Too close to call

Kimi K2.5 Instant: 40.2 (#123), Qwen3 14B: 39.3 (#134)

Knowledge benchmarks
BenchmarkKimi K2.5 InstantQwen3 14B
GPQA Diamond—63.8%
Vectara Hallucination Rate—5.4%
LMArena Expert1440—

Multimodal Not comparable

Kimi K2.5 Instant: 40.2 (#50), Qwen3 14B: —

Multimodal benchmarks
BenchmarkKimi K2.5 InstantQwen3 14B
LMArena Vision1254—

Multilingual Not comparable

Kimi K2.5 Instant: 52.0 (#94), Qwen3 14B: —

Multilingual benchmarks
BenchmarkKimi K2.5 InstantQwen3 14B
LMArena Non-English1406—
LMArena Chinese1449—
LMArena French1403—
LMArena German1413—
LMArena Korean1378—
LMArena Russian1404—
LMArena Spanish1447—

Instruction Following Not comparable

Kimi K2.5 Instant: 75.3 (#65), Qwen3 14B: —

Instruction Following benchmarks
BenchmarkKimi K2.5 InstantQwen3 14B
LMArena Instruction Following1430—

Long Context Kimi K2.5 Instant leads

Kimi K2.5 Instant: 43.9 (#83), Qwen3 14B: 38.1 (#204)

Long Context benchmarks
BenchmarkKimi K2.5 InstantQwen3 14B
Fiction.LiveBench—62.5%
LMArena Longer Query1435—

Writing & Preference Not comparable

Kimi K2.5 Instant: 60.6 (#95), Qwen3 14B: —

Writing & Preference benchmarks
BenchmarkKimi K2.5 InstantQwen3 14B
LMArena Text1420—
LMArena Creative Writing1381—
LMArena Multi-Turn1427—

Frequently asked questions

Is Kimi K2.5 Instant better than Qwen3 14B?

Kimi K2.5 Instant is the stronger model overall, scoring 43.6 to 35.5 on the Noometry Index.

Is Kimi K2.5 Instant or Qwen3 14B better for coding?

Kimi K2.5 Instant scores higher on coding benchmarks: 42.6 versus 37.3 in the Noometry coding category.

How many benchmarks do Kimi K2.5 Instant and Qwen3 14B share?

0 benchmarks have published results for both models. Kimi K2.5 Instant has 18 scored results on Noometry and Qwen3 14B has 12.

Related comparisons

Go deeper