Model comparison

Claude Sonnet 4 vs Kimi K2.5

Kimi K2.5 is the stronger model overall, scoring 48.1 to 40.8 on the Noometry Index.

Last verified . 38 shared benchmarks.

Claude Sonnet 4 Anthropic

40.8

Rank #145 Confirmed

Kimi K2.5 Moonshot AI

48.1

Rank #57 Confirmed

Summary

  • They share 38 benchmarks with published results for both. Claude Sonnet 4 scores higher in 1 category and Kimi K2.5 in 9 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Kimi K2.5 leads 52.1 to 33.7.
  • The biggest single-benchmark swing is Fiction.LiveBench: 46.9% for Claude Sonnet 4 and 86.1% for Kimi K2.5.
  • Kimi K2.5 is cheaper at $0.45 / $2.25 per million input/output tokens, against $3 / $15 for Claude Sonnet 4.
  • Kimi K2.5 accepts more context: 262K tokens versus 200K.
  • Kimi K2.5 has downloadable open weights; the other is API-only.

Side by side

Claude Sonnet 4 and Kimi K2.5 specifications
Claude Sonnet 4Kimi K2.5
ProviderAnthropicMoonshot AI
Noometry Index40.848.1
Released2025-05-222026-01-27
WeightsProprietaryOpen
Context window200K262K
Max output64K262K
Input $ / M tokens$3$0.45
Output $ / M tokens$15$2.25
Results tracked5851

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2.5 leads

Claude Sonnet 4: 43.5 (#88), Kimi K2.5: 48.8 (#53)

Coding benchmarks
BenchmarkClaude Sonnet 4Kimi K2.5
SWE-bench Verified (bash only)64.9%70.8%
SciCode40%49%
WeirdML46.1%45.6%
LMArena Coding14141474
ALE-Bench655.35821.65
SWE-bench Verified—73.8%
Aider Polyglot61.3%—
LMArena WebDev—1437
SWE-bench Multilingual—67.3%
GSO4.9%—

Agentic & Tool Use Claude Sonnet 4 leads

Claude Sonnet 4: 38.5 (#31), Kimi K2.5: 34.2 (#48)

Agentic & Tool Use benchmarks
BenchmarkClaude Sonnet 4Kimi K2.5
OSWorld43.9%63.3%
Terminal-Bench—43.2%
TheAgentCompany33.1%—
Cybench35%—
DeepResearch Bench46.6%—
METR Time Horizons62%—
Vending-Bench 2—1,198

Reasoning Kimi K2.5 leads

Claude Sonnet 4: 22.9 (#187), Kimi K2.5: 31.2 (#80)

Reasoning benchmarks
BenchmarkClaude Sonnet 4Kimi K2.5
ARC-AGI-25.9%11.8%
SimpleBench45.5%46.8%
Kagi LLM Benchmark73%78.5%
ARC-AGI-140%65.3%
CritPt0.3%3.1%
EnigmaEval3.1%3.4%
LMArena Hard Prompts13721453
Epoch Capabilities Index141.69148.03
NYT Connections (extended)—69.9%
Chess Puzzles—12%
Thematic Generalization—69.4%
DTBench77.1%—
LMCA29%—
ForecastBench60.2—

Math Kimi K2.5 leads

Claude Sonnet 4: 43.3 (#80), Kimi K2.5: 51.8 (#53)

Math benchmarks
BenchmarkClaude Sonnet 4Kimi K2.5
OTIS Mock AIME 2024-202571.1%92.2%
LMArena Math13751470
FrontierMath (Feb 2025 set)4.1%27.9%
FrontierMath Tier 4 (v1)0%4.2%
MathArena Final-Answer Competitions—62.3%
Omni-MATH60.2%—
MATH Level 584.4%—

Knowledge Kimi K2.5 leads

Claude Sonnet 4: 41.8 (#108), Kimi K2.5: 53.6 (#56)

Knowledge benchmarks
BenchmarkClaude Sonnet 4Kimi K2.5
GPQA Diamond79.2%87.6%
Humanity's Last Exam7.8%24.4%
Vectara Hallucination Rate10.3%14.2%
LMArena Expert13721466
SimpleQA Verified—34.3%
MMLU-Pro84.3%—
Confabulations13.2%—
GPQA (HELM)70.6%—

Multimodal Kimi K2.5 leads

Claude Sonnet 4: 26.2 (#121), Kimi K2.5: 41.1 (#39)

Multimodal benchmarks
BenchmarkClaude Sonnet 4Kimi K2.5
LMArena Vision11911269
GeoBench37%—
VPCT34%—
LMArena Document—1430
MindCube44.8%—

Multilingual Kimi K2.5 leads

Claude Sonnet 4: 46.7 (#156), Kimi K2.5: 53.9 (#53)

Multilingual benchmarks
BenchmarkClaude Sonnet 4Kimi K2.5
LMArena Non-English13331433
LMArena Chinese13501495
LMArena French13631454
LMArena German13311441
LMArena Japanese13021421
LMArena Korean12911410
LMArena Russian13551435
LMArena Spanish13571450

Instruction Following Kimi K2.5 leads

Claude Sonnet 4: 71.7 (#145), Kimi K2.5: 75.3 (#64)

Instruction Following benchmarks
BenchmarkClaude Sonnet 4Kimi K2.5
LMArena Instruction Following13761431
IFEval84%—

Long Context Kimi K2.5 leads

Claude Sonnet 4: 33.7 (#259), Kimi K2.5: 52.1 (#7)

Long Context benchmarks
BenchmarkClaude Sonnet 4Kimi K2.5
Fiction.LiveBench46.9%86.1%
LMArena Longer Query13981445
CL-bench—19.3%
CL-bench Life—13.2%

Writing & Preference Kimi K2.5 leads

Claude Sonnet 4: 57.1 (#132), Kimi K2.5: 65.1 (#53)

Writing & Preference benchmarks
BenchmarkClaude Sonnet 4Kimi K2.5
LMArena Text13511445
LMArena Creative Writing13451423
EQ-Bench Creative Writing14831579
LMArena Multi-Turn13761444
Short-Story Creative Writing81.4%—
WildBench83.8%—

Frequently asked questions

Is Claude Sonnet 4 better than Kimi K2.5?

Kimi K2.5 is the stronger model overall, scoring 48.1 to 40.8 on the Noometry Index.

Which is cheaper, Claude Sonnet 4 or Kimi K2.5?

Kimi K2.5 is cheaper. It lists at $0.45 per million input tokens and $2.25 per million output tokens; Claude Sonnet 4 lists at $3 and $15.

Is Claude Sonnet 4 or Kimi K2.5 better for coding?

Kimi K2.5 scores higher on coding benchmarks: 48.8 versus 43.5 in the Noometry coding category.

Which has the bigger context window?

Kimi K2.5 does, with 262K tokens against 200K.

How many benchmarks do Claude Sonnet 4 and Kimi K2.5 share?

38 benchmarks have published results for both models. Claude Sonnet 4 has 58 scored results on Noometry and Kimi K2.5 has 51.

Related comparisons

Go deeper