Model comparison

Kimi K2 (Jul 2025) vs Qwen3 Max

Qwen3 Max is the stronger model overall, scoring 43.7 to 41.2 on the Noometry Index. Kimi K2 (Jul 2025) costs 2.4× less per token, which makes it the better buy when Qwen3 Max's lead doesn't matter for your workload.

Last verified . 22 shared benchmarks.

Kimi K2 (Jul 2025) Moonshot AI

41.2

Rank #140 Confirmed

Qwen3 Max Alibaba (Qwen)

43.7

Rank #87 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Kimi K2 (Jul 2025) scores higher in 2 categories and Qwen3 Max in 6 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3 Max leads 48.1 to 37.3.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 64.4% for Kimi K2 (Jul 2025) and 72.5% for Qwen3 Max.
  • Kimi K2 (Jul 2025) is cheaper at $0.57 / $2.30 per million input/output tokens, against $1.20 / $6 for Qwen3 Max.
  • Kimi K2 (Jul 2025) has downloadable open weights; the other is API-only.

Side by side

Kimi K2 (Jul 2025) and Qwen3 Max specifications
Kimi K2 (Jul 2025)Qwen3 Max
ProviderMoonshot AIAlibaba (Qwen)
Noometry Index41.243.7
Released2025-07-122025-09-23
WeightsOpenProprietary
Context window262K262K
Max output262K66K
Input $ / M tokens$0.57$1.20
Output $ / M tokens$2.30$6
Results tracked4233

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Kimi K2 (Jul 2025): 42.4 (#102), Qwen3 Max: 43.0 (#93)

Coding benchmarks
BenchmarkKimi K2 (Jul 2025)Qwen3 Max
LMArena Coding13991456
ALE-Bench597.5370.45
SWE-bench Verified (bash only)63.4%—
Aider Polyglot59.1%—
GSO4.9%—
WeirdML42.8%—

Agentic & Tool Use Not comparable

Kimi K2 (Jul 2025): 32.4 (#64), Qwen3 Max: —

Agentic & Tool Use benchmarks
BenchmarkKimi K2 (Jul 2025)Qwen3 Max
Terminal-Bench35.7%—
Berkeley Function Calling Leaderboard59.1%—
METR Time Horizons59.2%—
Vending-Bench 2—71.56

Reasoning Too close to call

Kimi K2 (Jul 2025): 23.3 (#179), Qwen3 Max: 22.6 (#190)

Reasoning benchmarks
BenchmarkKimi K2 (Jul 2025)Qwen3 Max
Kagi LLM Benchmark64.4%72.5%
LMArena Hard Prompts13841448
Epoch Capabilities Index146.01142.38
SimpleBench26.3%—
NYT Connections (extended)—30.1%
Chess Puzzles—4%
Mystery Game Puzzles—5%
DTBench—82.1%
LMCA—28.3%
ForecastBench60.2—

Math Kimi K2 (Jul 2025) leads

Kimi K2 (Jul 2025): 42.7 (#83), Qwen3 Max: 38.7 (#131)

Math benchmarks
BenchmarkKimi K2 (Jul 2025)Qwen3 Max
LMArena Math13971446
FrontierMath (Tiers 1-3)—18.9%
OTIS Mock AIME 2024-2025—73.3%
Omni-MATH65.4%—
MATH Level 5—97.1%
FrontierMath (Feb 2025 set)21.4%—
FrontierMath Tier 4 (v1)0%—

Knowledge Qwen3 Max leads

Kimi K2 (Jul 2025): 37.3 (#157), Qwen3 Max: 48.1 (#78)

Knowledge benchmarks
BenchmarkKimi K2 (Jul 2025)Qwen3 Max
LMArena Expert13651455
GPQA Diamond—72.6%
SimpleQA Verified—48.7%
MMLU-Pro81.9%—
Confabulations20.4%—
Vectara Hallucination Rate17.9%—
GPQA (HELM)65.3%—

Multilingual Qwen3 Max leads

Kimi K2 (Jul 2025): 49.6 (#130), Qwen3 Max: 53.7 (#62)

Multilingual benchmarks
BenchmarkKimi K2 (Jul 2025)Qwen3 Max
LMArena Non-English13721429
LMArena Chinese14151478
LMArena French13791449
LMArena German13871463
LMArena Japanese13491397
LMArena Korean13251399
LMArena Russian13851428
LMArena Spanish13861462

Instruction Following Qwen3 Max leads

Kimi K2 (Jul 2025): 71.1 (#156), Qwen3 Max: 74.8 (#87)

Instruction Following benchmarks
BenchmarkKimi K2 (Jul 2025)Qwen3 Max
LMArena Instruction Following13481419
IFEval85%—

Long Context Too close to call

Kimi K2 (Jul 2025): 41.2 (#145), Qwen3 Max: 41.6 (#134)

Long Context benchmarks
BenchmarkKimi K2 (Jul 2025)Qwen3 Max
Fiction.LiveBench66.7%66.7%
CL-bench17.6%14.5%
LMArena Longer Query13531438

Writing & Preference Too close to call

Kimi K2 (Jul 2025): 62.3 (#78), Qwen3 Max: 62.4 (#76)

Writing & Preference benchmarks
BenchmarkKimi K2 (Jul 2025)Qwen3 Max
LMArena Text13801439
LMArena Creative Writing13501402
LMArena Multi-Turn13711446
Short-Story Creative Writing85.6%—
EQ-Bench Creative Writing1666—
WildBench86.2%—

Frequently asked questions

Is Kimi K2 (Jul 2025) better than Qwen3 Max?

Qwen3 Max is the stronger model overall, scoring 43.7 to 41.2 on the Noometry Index. Kimi K2 (Jul 2025) costs 2.4× less per token, which makes it the better buy when Qwen3 Max's lead doesn't matter for your workload.

Which is cheaper, Kimi K2 (Jul 2025) or Qwen3 Max?

Kimi K2 (Jul 2025) is cheaper. It lists at $0.57 per million input tokens and $2.30 per million output tokens; Qwen3 Max lists at $1.20 and $6.

Is Kimi K2 (Jul 2025) or Qwen3 Max better for coding?

They score almost the same on coding (42.4 vs 43.0); test both on your own repository before choosing.

Which has the bigger context window?

Both accept 262K tokens.

How many benchmarks do Kimi K2 (Jul 2025) and Qwen3 Max share?

22 benchmarks have published results for both models. Kimi K2 (Jul 2025) has 42 scored results on Noometry and Qwen3 Max has 33.

Related comparisons

Go deeper