Model comparison

Kimi K2.7 Code vs Qwen3.5 397B-A17B

Qwen3.5 397B-A17B is the stronger model overall, scoring 46.0 to 43.3 on the Noometry Index.

Last verified . 7 shared benchmarks.

Kimi K2.7 Code Moonshot AI

43.3

Rank #94 Confirmed

Qwen3.5 397B-A17B Alibaba (Qwen)

46.0

Rank #67 Confirmed

Summary

  • They share 7 benchmarks with published results for both. Kimi K2.7 Code scores higher in 4 categories and Qwen3.5 397B-A17B in 1 category; 3 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where Qwen3.5 397B-A17B leads 33.3 to 24.0.
  • The biggest single-benchmark swing is FrontierMath (Tiers 1-3): 54% for Kimi K2.7 Code and 31.2% for Qwen3.5 397B-A17B.
  • Qwen3.5 397B-A17B is cheaper at $0.60 / $3.60 per million input/output tokens, against $0.95 / $4 for Kimi K2.7 Code.

Side by side

Kimi K2.7 Code and Qwen3.5 397B-A17B specifications
Kimi K2.7 CodeQwen3.5 397B-A17B
ProviderMoonshot AIAlibaba (Qwen)
Noometry Index43.346.0
Released2026-06-122026-02-01
WeightsOpenOpen
Context window262K262K
Max output262K66K
Input $ / M tokens$0.95$0.60
Output $ / M tokens$4$3.60
Results tracked1936

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Kimi K2.7 Code: 42.9 (#95), Qwen3.5 397B-A17B: 42.0 (#114)

Coding benchmarks
BenchmarkKimi K2.7 CodeQwen3.5 397B-A17B
LMArena WebDev14731400
DeepSWE30.5%—
FrontierCode30.1%—
SciCode47.5%—
WeirdML54.1%—
LMArena Coding—1465
ALE-Bench886.23—

Agentic & Tool Use Qwen3.5 397B-A17B leads

Kimi K2.7 Code: 24.0 (#122), Qwen3.5 397B-A17B: 33.3 (#53)

Agentic & Tool Use benchmarks
BenchmarkKimi K2.7 CodeQwen3.5 397B-A17B
APEX-Agents37.6%24.9%
τ²-bench Airline—81.5%
τ²-bench Banking—9.8%
τ²-bench Retail—84.4%
τ²-bench Telecom—97.8%
GBAEval0.9%—
Vending-Bench 25,083—

Reasoning Kimi K2.7 Code leads

Kimi K2.7 Code: 39.0 (#61), Qwen3.5 397B-A17B: 34.5 (#70)

Reasoning benchmarks
BenchmarkKimi K2.7 CodeQwen3.5 397B-A17B
Chess Puzzles21%13%
Epoch Capabilities Index149.97146.65
SimpleBench57.9%—
Kagi LLM Benchmark—73.7%
NYT Connections (extended)—58.9%
CritPt10%—
Thematic Generalization—65.1%
LMArena Hard Prompts—1448
Mystery Game Puzzles—18%
DTBench—87.5%
LMCA—37.9%
Surface Evolver Bench48.8%—

Math Kimi K2.7 Code leads

Kimi K2.7 Code: 52.9 (#48), Qwen3.5 397B-A17B: 46.1 (#73)

Math benchmarks
BenchmarkKimi K2.7 CodeQwen3.5 397B-A17B
FrontierMath (Tiers 1-3)54%31.2%
OTIS Mock AIME 2024-202595.6%88.9%
FrontierMath Tier 412.2%—
LMArena Math—1454

Knowledge Too close to call

Kimi K2.7 Code: 53.5 (#57), Qwen3.5 397B-A17B: 53.3 (#58)

Knowledge benchmarks
BenchmarkKimi K2.7 CodeQwen3.5 397B-A17B
GPQA Diamond87.9%86.4%
SimpleQA Verified36.5%—
LMArena Expert—1462

Multimodal Not comparable

Kimi K2.7 Code: —, Qwen3.5 397B-A17B: 40.7 (#44)

Multimodal benchmarks
BenchmarkKimi K2.7 CodeQwen3.5 397B-A17B
LMArena Vision—1263

Multilingual Not comparable

Kimi K2.7 Code: —, Qwen3.5 397B-A17B: 53.7 (#59)

Multilingual benchmarks
BenchmarkKimi K2.7 CodeQwen3.5 397B-A17B
LMArena Non-English—1430
LMArena Chinese—1500
LMArena French—1461
LMArena German—1447
LMArena Japanese—1426
LMArena Korean—1384
LMArena Russian—1429
LMArena Spanish—1441

Instruction Following Not comparable

Kimi K2.7 Code: —, Qwen3.5 397B-A17B: 75.0 (#77)

Instruction Following benchmarks
BenchmarkKimi K2.7 CodeQwen3.5 397B-A17B
LMArena Instruction Following—1424

Long Context Not comparable

Kimi K2.7 Code: —, Qwen3.5 397B-A17B: 44.1 (#74)

Long Context benchmarks
BenchmarkKimi K2.7 CodeQwen3.5 397B-A17B
LMArena Longer Query—1442

Writing & Preference Not comparable

Kimi K2.7 Code: —, Qwen3.5 397B-A17B: 62.3 (#79)

Writing & Preference benchmarks
BenchmarkKimi K2.7 CodeQwen3.5 397B-A17B
LMArena Text—1438
LMArena Creative Writing—1401
EQ-Bench Creative Writing—1478
LMArena Multi-Turn—1446

Frequently asked questions

Is Kimi K2.7 Code better than Qwen3.5 397B-A17B?

Qwen3.5 397B-A17B is the stronger model overall, scoring 46.0 to 43.3 on the Noometry Index.

Which is cheaper, Kimi K2.7 Code or Qwen3.5 397B-A17B?

Qwen3.5 397B-A17B is cheaper. It lists at $0.60 per million input tokens and $3.60 per million output tokens; Kimi K2.7 Code lists at $0.95 and $4.

Is Kimi K2.7 Code or Qwen3.5 397B-A17B better for coding?

They score almost the same on coding (42.9 vs 42.0); test both on your own repository before choosing.

Which has the bigger context window?

Both accept 262K tokens.

How many benchmarks do Kimi K2.7 Code and Qwen3.5 397B-A17B share?

7 benchmarks have published results for both models. Kimi K2.7 Code has 19 scored results on Noometry and Qwen3.5 397B-A17B has 36.

Related comparisons

Go deeper