Model comparison

Claude 3.5 Sonnet vs Kimi K2 (Jul 2025)

Kimi K2 (Jul 2025) is the stronger model overall, scoring 41.2 to 34.6 on the Noometry Index.

Last verified . 34 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

Kimi K2 (Jul 2025) Moonshot AI

41.2

Rank #140 Confirmed

Summary

  • They share 34 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 0 categories and Kimi K2 (Jul 2025) in 9 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Kimi K2 (Jul 2025) leads 42.7 to 19.2.
  • The biggest single-benchmark swing is Omni-MATH: 27.6% for Claude 3.5 Sonnet and 65.4% for Kimi K2 (Jul 2025).
  • Kimi K2 (Jul 2025) has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Sonnet and Kimi K2 (Jul 2025) specifications
Claude 3.5 SonnetKimi K2 (Jul 2025)
ProviderAnthropicMoonshot AI
Noometry Index34.641.2
Released2024-06-202025-07-12
WeightsProprietaryOpen
Context window—262K
Max output—262K
Input $ / M tokens—$0.57
Output $ / M tokens—$2.30
Results tracked6042

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2 (Jul 2025) leads

Claude 3.5 Sonnet: 39.0 (#165), Kimi K2 (Jul 2025): 42.4 (#102)

Coding benchmarks
BenchmarkClaude 3.5 SonnetKimi K2 (Jul 2025)
Aider Polyglot51.6%59.1%
GSO4.6%4.9%
WeirdML40%42.8%
LMArena Coding13421399
SWE-bench Verified (bash only)—63.4%
BigCodeBench Instruct46.8%—
LiveBench Coding67.1%—
BigCodeBench Complete58.6%—
CadEval48%—
ALE-Bench—597.5
HumanEval+81.7%—
MBPP+74.3%—

Agentic & Tool Use Too close to call

Claude 3.5 Sonnet: 32.3 (#67), Kimi K2 (Jul 2025): 32.4 (#64)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 SonnetKimi K2 (Jul 2025)
METR Time Horizons45.2%59.2%
Terminal-Bench—35.7%
Berkeley Function Calling Leaderboard—59.1%
TheAgentCompany24%—
Cybench17.5%—
BALROG32.6%—

Reasoning Too close to call

Claude 3.5 Sonnet: 23.1 (#183), Kimi K2 (Jul 2025): 23.3 (#179)

Reasoning benchmarks
BenchmarkClaude 3.5 SonnetKimi K2 (Jul 2025)
SimpleBench41.4%26.3%
LMArena Hard Prompts13051384
Epoch Capabilities Index133.55146.01
ForecastBench60.760.2
Kagi LLM Benchmark—64.4%
EnigmaEval0.9%—
LiveBench Reasoning56.7%—
DTBench67.8%—
LiveBench Data Analysis55%—
LiveBench59%—

Math Kimi K2 (Jul 2025) leads

Claude 3.5 Sonnet: 19.2 (#288), Kimi K2 (Jul 2025): 42.7 (#83)

Math benchmarks
BenchmarkClaude 3.5 SonnetKimi K2 (Jul 2025)
Omni-MATH27.6%65.4%
LMArena Math13071397
FrontierMath (Feb 2025 set)2.1%21.4%
FrontierMath Tier 4 (v1)0%0%
OTIS Mock AIME 2024-20258.5%—
LiveBench Math52.3%—
MATH Level 556.9%—

Knowledge Kimi K2 (Jul 2025) leads

Claude 3.5 Sonnet: 28.6 (#245), Kimi K2 (Jul 2025): 37.3 (#157)

Knowledge benchmarks
BenchmarkClaude 3.5 SonnetKimi K2 (Jul 2025)
MMLU-Pro77.7%81.9%
Confabulations19.9%20.4%
GPQA (HELM)56.5%65.3%
LMArena Expert12651365
GPQA Diamond55.3%—
Humanity's Last Exam4.1%—
Vectara Hallucination Rate—17.9%
MMLU87.3%—

Multimodal Not comparable

Claude 3.5 Sonnet: 26.5 (#120), Kimi K2 (Jul 2025): —

Multimodal benchmarks
BenchmarkClaude 3.5 SonnetKimi K2 (Jul 2025)
LMArena Vision1125—
Video-MME60%—
GeoBench62%—
VPCT33%—

Multilingual Kimi K2 (Jul 2025) leads

Claude 3.5 Sonnet: 43.2 (#185), Kimi K2 (Jul 2025): 49.6 (#130)

Multilingual benchmarks
BenchmarkClaude 3.5 SonnetKimi K2 (Jul 2025)
LMArena Non-English12831372
LMArena Chinese12721415
LMArena French13051379
LMArena German12971387
LMArena Japanese12341349
LMArena Korean12001325
LMArena Russian13061385
LMArena Spanish12901386

Instruction Following Kimi K2 (Jul 2025) leads

Claude 3.5 Sonnet: 68.8 (#182), Kimi K2 (Jul 2025): 71.1 (#156)

Instruction Following benchmarks
BenchmarkClaude 3.5 SonnetKimi K2 (Jul 2025)
IFEval85.5%85%
LMArena Instruction Following12971348
LiveBench Instruction Following69.3%—

Long Context Kimi K2 (Jul 2025) leads

Claude 3.5 Sonnet: 39.9 (#167), Kimi K2 (Jul 2025): 41.2 (#145)

Long Context benchmarks
BenchmarkClaude 3.5 SonnetKimi K2 (Jul 2025)
LMArena Longer Query13111353
Fiction.LiveBench—66.7%
CL-bench—17.6%

Writing & Preference Kimi K2 (Jul 2025) leads

Claude 3.5 Sonnet: 52.9 (#164), Kimi K2 (Jul 2025): 62.3 (#78)

Writing & Preference benchmarks
BenchmarkClaude 3.5 SonnetKimi K2 (Jul 2025)
LMArena Text12981380
LMArena Creative Writing12921350
Short-Story Creative Writing80.3%85.6%
EQ-Bench Creative Writing14511666
WildBench79.2%86.2%
LMArena Multi-Turn13261371
LiveBench Language53.8%—

Frequently asked questions

Is Claude 3.5 Sonnet better than Kimi K2 (Jul 2025)?

Kimi K2 (Jul 2025) is the stronger model overall, scoring 41.2 to 34.6 on the Noometry Index.

Is Claude 3.5 Sonnet or Kimi K2 (Jul 2025) better for coding?

Kimi K2 (Jul 2025) scores higher on coding benchmarks: 42.4 versus 39.0 in the Noometry coding category.

How many benchmarks do Claude 3.5 Sonnet and Kimi K2 (Jul 2025) share?

34 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and Kimi K2 (Jul 2025) has 42.

Related comparisons

Go deeper