Model comparison

Claude Sonnet 4 vs Kimi K2.6

Kimi K2.6 is the stronger model overall, scoring 47.7 to 40.8 on the Noometry Index.

Last verified . 31 shared benchmarks.

Claude Sonnet 4 Anthropic

40.8

Rank #145 Confirmed

Kimi K2.6 Moonshot AI

47.7

Rank #60 Confirmed

Summary

  • They share 31 benchmarks with published results for both. Claude Sonnet 4 scores higher in 1 category and Kimi K2.6 in 9 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Kimi K2.6 leads 40.5 to 22.9.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 71.1% for Claude Sonnet 4 and 96.1% for Kimi K2.6.
  • Kimi K2.6 is cheaper at $0.95 / $4 per million input/output tokens, against $3 / $15 for Claude Sonnet 4.
  • Kimi K2.6 accepts more context: 262K tokens versus 200K.
  • Kimi K2.6 has downloadable open weights; the other is API-only.

Side by side

Claude Sonnet 4 and Kimi K2.6 specifications
Claude Sonnet 4Kimi K2.6
ProviderAnthropicMoonshot AI
Noometry Index40.847.7
Released2025-05-222026-04-20
WeightsProprietaryOpen
Context window200K262K
Max output64K262K
Input $ / M tokens$3$0.95
Output $ / M tokens$15$4
Results tracked5851

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2.6 leads

Claude Sonnet 4: 43.5 (#88), Kimi K2.6: 50.7 (#43)

Coding benchmarks
BenchmarkClaude Sonnet 4Kimi K2.6
SciCode40%53.5%
WeirdML46.1%55.9%
LMArena Coding14141488
ALE-Bench655.351,093
SWE-bench Verified—76.7%
SWE-bench Verified (bash only)64.9%—
Aider Polyglot61.3%—
LMArena WebDev—1509
GSO4.9%—

Agentic & Tool Use Claude Sonnet 4 leads

Claude Sonnet 4: 38.5 (#31), Kimi K2.6: 21.9 (#137)

Agentic & Tool Use benchmarks
BenchmarkClaude Sonnet 4Kimi K2.6
OSWorld 2.0—4.6%
TheAgentCompany33.1%—
Cybench35%—
DeepResearch Bench46.6%—
OSWorld43.9%—
ExploitBench—18.4%
GBAEval—0.9%
GDP.pdf—12%
METR Time Horizons62%—
Vending-Bench 2—6,205

Reasoning Kimi K2.6 leads

Claude Sonnet 4: 22.9 (#187), Kimi K2.6: 40.5 (#55)

Reasoning benchmarks
BenchmarkClaude Sonnet 4Kimi K2.6
CritPt0.3%8%
LMArena Hard Prompts13721470
DTBench77.1%90.9%
LMCA29%37.3%
Epoch Capabilities Index141.69151.05
ARC-AGI-25.9%—
SimpleBench45.5%—
Kagi LLM Benchmark73%—
NYT Connections (extended)—87.2%
ARC-AGI-140%—
Chess Puzzles—26%
EnigmaEval3.1%—
EBR-Bench—2.4%
Mystery Game Puzzles—18%
ForecastBench60.2—

Math Kimi K2.6 leads

Claude Sonnet 4: 43.3 (#80), Kimi K2.6: 57.0 (#41)

Knowledge Kimi K2.6 leads

Claude Sonnet 4: 41.8 (#108), Kimi K2.6: 54.0 (#54)

Knowledge benchmarks
BenchmarkClaude Sonnet 4Kimi K2.6
GPQA Diamond79.2%90.8%
Vectara Hallucination Rate10.3%10.8%
LMArena Expert13721491
Humanity's Last Exam7.8%—
SimpleQA Verified—34.9%
MMLU-Pro84.3%—
Confabulations13.2%—
GPQA (HELM)70.6%—

Multimodal Kimi K2.6 leads

Claude Sonnet 4: 26.2 (#121), Kimi K2.6: 31.6 (#103)

Multimodal benchmarks
BenchmarkClaude Sonnet 4Kimi K2.6
LMArena Vision11911283
GeoBench37%—
VPCT34%—
Blueprint-Bench 2—3.9%
Furniture Assembly—21.7%
LMArena Document—1451
MindCube44.8%—

Multilingual Kimi K2.6 leads

Claude Sonnet 4: 46.7 (#156), Kimi K2.6: 54.9 (#37)

Multilingual benchmarks
BenchmarkClaude Sonnet 4Kimi K2.6
LMArena Non-English13331446
LMArena Chinese13501521
LMArena French13631471
LMArena German13311450
LMArena Japanese13021443
LMArena Korean12911427
LMArena Russian13551446
LMArena Spanish13571464

Instruction Following Kimi K2.6 leads

Claude Sonnet 4: 71.7 (#145), Kimi K2.6: 76.3 (#43)

Instruction Following benchmarks
BenchmarkClaude Sonnet 4Kimi K2.6
LMArena Instruction Following13761451
IFEval84%—

Long Context Kimi K2.6 leads

Claude Sonnet 4: 33.7 (#259), Kimi K2.6: 44.9 (#52)

Long Context benchmarks
BenchmarkClaude Sonnet 4Kimi K2.6
LMArena Longer Query13981468
Fiction.LiveBench46.9%—

Writing & Preference Kimi K2.6 leads

Claude Sonnet 4: 57.1 (#132), Kimi K2.6: 68.5 (#26)

Writing & Preference benchmarks
BenchmarkClaude Sonnet 4Kimi K2.6
LMArena Text13511455
LMArena Creative Writing13451434
EQ-Bench Creative Writing14831725
LMArena Multi-Turn13761453
Short-Story Creative Writing81.4%—
WildBench83.8%—
EQ-Bench 4—1202

Frequently asked questions

Is Claude Sonnet 4 better than Kimi K2.6?

Kimi K2.6 is the stronger model overall, scoring 47.7 to 40.8 on the Noometry Index.

Which is cheaper, Claude Sonnet 4 or Kimi K2.6?

Kimi K2.6 is cheaper. It lists at $0.95 per million input tokens and $4 per million output tokens; Claude Sonnet 4 lists at $3 and $15.

Is Claude Sonnet 4 or Kimi K2.6 better for coding?

Kimi K2.6 scores higher on coding benchmarks: 50.7 versus 43.5 in the Noometry coding category.

Which has the bigger context window?

Kimi K2.6 does, with 262K tokens against 200K.

How many benchmarks do Claude Sonnet 4 and Kimi K2.6 share?

31 benchmarks have published results for both models. Claude Sonnet 4 has 58 scored results on Noometry and Kimi K2.6 has 51.

Related comparisons

Go deeper