Model comparison

Claude Sonnet 5 vs Kimi K3

Kimi K3 is the stronger model overall, scoring 59.5 to 54.6 on the Noometry Index. Claude Sonnet 5 costs 1.5× less per token, which makes it the better buy when Kimi K3's lead doesn't matter for your workload.

Last verified . 45 shared benchmarks.

Claude Sonnet 5 Anthropic

54.6

Rank #29 Confirmed

Kimi K3 Moonshot AI

59.5

Rank #15 Confirmed

Summary

  • They share 45 benchmarks with published results for both. Claude Sonnet 5 scores higher in 2 categories and Kimi K3 in 8 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Kimi K3 leads 63.0 to 49.1.
  • The biggest single-benchmark swing is Surface Evolver Bench: 60% for Claude Sonnet 5 and 95% for Kimi K3.
  • Claude Sonnet 5 is cheaper at $2 / $10 per million input/output tokens, against $3 / $15 for Kimi K3.
  • Kimi K3 accepts more context: 1.05M tokens versus 1M.
  • Kimi K3 has downloadable open weights; the other is API-only.

Side by side

Claude Sonnet 5 and Kimi K3 specifications
Claude Sonnet 5Kimi K3
ProviderAnthropicMoonshot AI
Noometry Index54.659.5
Released2026-06-292026-07-16
WeightsProprietaryOpen
Context window1M1.05M
Max output128K1.05M
Input $ / M tokens$2$3
Output $ / M tokens$10$15
Results tracked5153

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K3 leads

Claude Sonnet 5: 55.5 (#26), Kimi K3: 61.0 (#10)

Coding benchmarks
BenchmarkClaude Sonnet 5Kimi K3
DeepSWE53.8%68.5%
FrontierCode42.7%44.2%
LMArena WebDev15411654
SciCode54.3%59.5%
WeirdML68.8%82.6%
LMArena Coding14831508
ALE-Bench1,4631,524
CursorBench34.1%—
FrontierSWE—25.9%
GSO37.3%—

Agentic & Tool Use Too close to call

Claude Sonnet 5: 42.8 (#18), Kimi K3: 41.8 (#20)

Agentic & Tool Use benchmarks
BenchmarkClaude Sonnet 5Kimi K3
APEX-Agents54.5%50.6%
GBAEval65.3%48.3%
Vending-Bench 26,3785,165
τ²-bench Banking—37.1%
PostTrainBench—32%
GDP.pdf—19%
LMArena Search1194—

Reasoning Kimi K3 leads

Claude Sonnet 5: 49.1 (#39), Kimi K3: 63.0 (#17)

Reasoning benchmarks
BenchmarkClaude Sonnet 5Kimi K3
SimpleBench60.6%60.7%
NYT Connections (extended)75.1%93.6%
CritPt16.9%23.4%
Chess Puzzles35%39%
LMArena Hard Prompts14611496
Mystery Game Puzzles35%26%
DTBench92.5%91.2%
LMCA50%52.7%
Surface Evolver Bench60%95%
Epoch Capabilities Index156.21157.45
ForecastBench61.161.1
ARC-AGI-2—60.4%
ARC-AGI-1—94.5%
Bench to the Future 30.14—

Math Kimi K3 leads

Claude Sonnet 5: 66.2 (#27), Kimi K3: 74.2 (#16)

Math benchmarks
BenchmarkClaude Sonnet 5Kimi K3
FrontierMath (Tiers 1-3)65.6%72.2%
FrontierMath Tier 429.3%39%
OTIS Mock AIME 2024-202594.7%97.2%
ProofBench77%87%
LMArena Math14671491
MathArena Final-Answer Competitions—87.8%

Knowledge Kimi K3 leads

Claude Sonnet 5: 55.6 (#47), Kimi K3: 63.2 (#21)

Knowledge benchmarks
BenchmarkClaude Sonnet 5Kimi K3
GPQA Diamond90.5%93.1%
SimpleQA Verified33.7%50.6%
LMArena Expert14901521

Multimodal Claude Sonnet 5 leads

Claude Sonnet 5: 42.4 (#31), Kimi K3: 37.8 (#70)

Multimodal benchmarks
BenchmarkClaude Sonnet 5Kimi K3
Blueprint-Bench 224.9%29.5%
LMArena Vision1274—
Furniture Assembly—34.2%
LMArena Document1466—

Multilingual Kimi K3 leads

Claude Sonnet 5: 53.8 (#55), Kimi K3: 56.3 (#21)

Multilingual benchmarks
BenchmarkClaude Sonnet 5Kimi K3
LMArena Non-English14311466
LMArena Chinese14771529
LMArena French14601491
LMArena German14401488
LMArena Japanese14221487
LMArena Korean14111458
LMArena Russian14511482
LMArena Spanish14371472

Instruction Following Kimi K3 leads

Claude Sonnet 5: 76.3 (#41), Kimi K3: 77.7 (#14)

Instruction Following benchmarks
BenchmarkClaude Sonnet 5Kimi K3
LMArena Instruction Following14521483

Long Context Kimi K3 leads

Claude Sonnet 5: 44.8 (#55), Kimi K3: 45.8 (#29)

Long Context benchmarks
BenchmarkClaude Sonnet 5Kimi K3
LMArena Longer Query14631494

Writing & Preference Kimi K3 leads

Claude Sonnet 5: 69.2 (#25), Kimi K3: 76.6 (#4)

Writing & Preference benchmarks
BenchmarkClaude Sonnet 5Kimi K3
LMArena Text14421476
LMArena Creative Writing14161454
EQ-Bench Creative Writing17942082
EQ-Bench 412361339
LMArena Multi-Turn14541488

Frequently asked questions

Is Claude Sonnet 5 better than Kimi K3?

Kimi K3 is the stronger model overall, scoring 59.5 to 54.6 on the Noometry Index. Claude Sonnet 5 costs 1.5× less per token, which makes it the better buy when Kimi K3's lead doesn't matter for your workload.

Which is cheaper, Claude Sonnet 5 or Kimi K3?

Claude Sonnet 5 is cheaper. It lists at $2 per million input tokens and $10 per million output tokens; Kimi K3 lists at $3 and $15.

Is Claude Sonnet 5 or Kimi K3 better for coding?

Kimi K3 scores higher on coding benchmarks: 61.0 versus 55.5 in the Noometry coding category.

Which has the bigger context window?

Kimi K3 does, with 1.05M tokens against 1M.

How many benchmarks do Claude Sonnet 5 and Kimi K3 share?

45 benchmarks have published results for both models. Claude Sonnet 5 has 51 scored results on Noometry and Kimi K3 has 53.

Related comparisons

Go deeper