Model comparison

Kimi K2 (Jul 2025) vs Qwen3.5-Flash

Qwen3.5-Flash is the stronger model overall, scoring 42.5 to 41.2 on the Noometry Index.

Last verified . 22 shared benchmarks.

Kimi K2 (Jul 2025) Moonshot AI

41.2

Rank #140 Confirmed

Qwen3.5-Flash Alibaba (Qwen)

42.5

Rank #112 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Kimi K2 (Jul 2025) scores higher in 3 categories and Qwen3.5-Flash in 5 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen3.5-Flash leads 33.7 to 23.3.
  • The biggest single-benchmark swing is Vectara Hallucination Rate: 17.9% for Kimi K2 (Jul 2025) and 10.5% for Qwen3.5-Flash.
  • Qwen3.5-Flash is cheaper at $0.10 / $0.40 per million input/output tokens, against $0.57 / $2.30 for Kimi K2 (Jul 2025).
  • Qwen3.5-Flash accepts more context: 1M tokens versus 262K.
  • Kimi K2 (Jul 2025) has downloadable open weights; the other is API-only.

Side by side

Kimi K2 (Jul 2025) and Qwen3.5-Flash specifications
Kimi K2 (Jul 2025)Qwen3.5-Flash
ProviderMoonshot AIAlibaba (Qwen)
Noometry Index41.242.5
Released2025-07-122026-02-23
WeightsOpenProprietary
Context window262K1M
Max output262K66K
Input $ / M tokens$0.57$0.10
Output $ / M tokens$2.30$0.40
Results tracked4232

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2 (Jul 2025) leads

Kimi K2 (Jul 2025): 42.4 (#102), Qwen3.5-Flash: 34.2 (#242)

Coding benchmarks
BenchmarkKimi K2 (Jul 2025)Qwen3.5-Flash
LMArena Coding13991412
ALE-Bench597.5221.8
SWE-bench Verified (bash only)63.4%—
Aider Polyglot59.1%—
LMArena WebDev—1244
GSO4.9%—
WeirdML42.8%—

Agentic & Tool Use Not comparable

Kimi K2 (Jul 2025): 32.4 (#64), Qwen3.5-Flash: —

Agentic & Tool Use benchmarks
BenchmarkKimi K2 (Jul 2025)Qwen3.5-Flash
Terminal-Bench35.7%—
Berkeley Function Calling Leaderboard59.1%—
METR Time Horizons59.2%—
Vending-Bench 2—462.69

Reasoning Qwen3.5-Flash leads

Kimi K2 (Jul 2025): 23.3 (#179), Qwen3.5-Flash: 33.7 (#72)

Reasoning benchmarks
BenchmarkKimi K2 (Jul 2025)Qwen3.5-Flash
LMArena Hard Prompts13841403
Epoch Capabilities Index146.01143.98
SimpleBench26.3%—
Kagi LLM Benchmark64.4%—
Chess Puzzles—21%
Mystery Game Puzzles—20%
DTBench—82.9%
LMCA—29.1%
ForecastBench60.2—

Math Kimi K2 (Jul 2025) leads

Kimi K2 (Jul 2025): 42.7 (#83), Qwen3.5-Flash: 37.4 (#158)

Math benchmarks
BenchmarkKimi K2 (Jul 2025)Qwen3.5-Flash
LMArena Math13971407
FrontierMath (Feb 2025 set)21.4%6.2%
FrontierMath Tier 4 (v1)0%0%
FrontierMath (Tiers 1-3)—18.2%
OTIS Mock AIME 2024-2025—84.4%
Omni-MATH65.4%—

Knowledge Qwen3.5-Flash leads

Kimi K2 (Jul 2025): 37.3 (#157), Qwen3.5-Flash: 43.2 (#93)

Knowledge benchmarks
BenchmarkKimi K2 (Jul 2025)Qwen3.5-Flash
Vectara Hallucination Rate17.9%10.5%
LMArena Expert13651407
GPQA Diamond—82.3%
SimpleQA Verified—20.3%
MMLU-Pro81.9%—
Confabulations20.4%—
GPQA (HELM)65.3%—

Multilingual Too close to call

Kimi K2 (Jul 2025): 49.6 (#130), Qwen3.5-Flash: 50.5 (#121)

Multilingual benchmarks
BenchmarkKimi K2 (Jul 2025)Qwen3.5-Flash
LMArena Non-English13721385
LMArena Chinese14151446
LMArena French13791412
LMArena German13871390
LMArena Japanese13491368
LMArena Korean13251344
LMArena Russian13851379
LMArena Spanish13861400

Instruction Following Qwen3.5-Flash leads

Kimi K2 (Jul 2025): 71.1 (#156), Qwen3.5-Flash: 72.6 (#139)

Instruction Following benchmarks
BenchmarkKimi K2 (Jul 2025)Qwen3.5-Flash
LMArena Instruction Following13481374
IFEval85%—

Long Context Qwen3.5-Flash leads

Kimi K2 (Jul 2025): 41.2 (#145), Qwen3.5-Flash: 42.4 (#124)

Long Context benchmarks
BenchmarkKimi K2 (Jul 2025)Qwen3.5-Flash
LMArena Longer Query13531392
Fiction.LiveBench66.7%—
CL-bench17.6%—

Writing & Preference Kimi K2 (Jul 2025) leads

Kimi K2 (Jul 2025): 62.3 (#78), Qwen3.5-Flash: 57.9 (#122)

Writing & Preference benchmarks
BenchmarkKimi K2 (Jul 2025)Qwen3.5-Flash
LMArena Text13801397
LMArena Creative Writing13501343
LMArena Multi-Turn13711393
Short-Story Creative Writing85.6%—
EQ-Bench Creative Writing1666—
WildBench86.2%—

Frequently asked questions

Is Kimi K2 (Jul 2025) better than Qwen3.5-Flash?

Qwen3.5-Flash is the stronger model overall, scoring 42.5 to 41.2 on the Noometry Index.

Which is cheaper, Kimi K2 (Jul 2025) or Qwen3.5-Flash?

Qwen3.5-Flash is cheaper. It lists at $0.10 per million input tokens and $0.40 per million output tokens; Kimi K2 (Jul 2025) lists at $0.57 and $2.30.

Is Kimi K2 (Jul 2025) or Qwen3.5-Flash better for coding?

Kimi K2 (Jul 2025) scores higher on coding benchmarks: 42.4 versus 34.2 in the Noometry coding category.

Which has the bigger context window?

Qwen3.5-Flash does, with 1M tokens against 262K.

How many benchmarks do Kimi K2 (Jul 2025) and Qwen3.5-Flash share?

22 benchmarks have published results for both models. Kimi K2 (Jul 2025) has 42 scored results on Noometry and Qwen3.5-Flash has 32.

Related comparisons

Go deeper