Model comparison

Claude 3.7 Sonnet vs Kimi K2 (Jul 2025)

Kimi K2 (Jul 2025) is the stronger model overall, scoring 41.2 to 39.5 on the Noometry Index.

Last verified . 34 shared benchmarks.

Claude 3.7 Sonnet Anthropic

39.5

Rank #164 Confirmed

Kimi K2 (Jul 2025) Moonshot AI

41.2

Rank #140 Confirmed

Summary

  • They share 34 benchmarks with published results for both. Claude 3.7 Sonnet scores higher in 4 categories and Kimi K2 (Jul 2025) in 5 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Claude 3.7 Sonnet leads 50.3 to 41.2.
  • The biggest single-benchmark swing is Omni-MATH: 33% for Claude 3.7 Sonnet and 65.4% for Kimi K2 (Jul 2025).
  • Kimi K2 (Jul 2025) has downloadable open weights; the other is API-only.

Side by side

Claude 3.7 Sonnet and Kimi K2 (Jul 2025) specifications
Claude 3.7 SonnetKimi K2 (Jul 2025)
ProviderAnthropicMoonshot AI
Noometry Index39.541.2
Released2025-02-242025-07-12
WeightsProprietaryOpen
Context window—262K
Max output—262K
Input $ / M tokens—$0.57
Output $ / M tokens—$2.30
Results tracked5842

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2 (Jul 2025) leads

Claude 3.7 Sonnet: 40.6 (#136), Kimi K2 (Jul 2025): 42.4 (#102)

Coding benchmarks
BenchmarkClaude 3.7 SonnetKimi K2 (Jul 2025)
SWE-bench Verified (bash only)52.8%63.4%
Aider Polyglot64.9%59.1%
GSO3.8%4.9%
LMArena Coding13611399
SWE-bench Verified61%—
WeirdML—42.8%
LiveBench Coding74.5%—
CadEval54%—
ALE-Bench—597.5

Agentic & Tool Use Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 34.1 (#50), Kimi K2 (Jul 2025): 32.4 (#64)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.7 SonnetKimi K2 (Jul 2025)
METR Time Horizons60%59.2%
Terminal-Bench—35.7%
Berkeley Function Calling Leaderboard—59.1%
TheAgentCompany30.9%—
Cybench20%—
DeepResearch Bench43.6%—
OSWorld35.8%—

Reasoning Kimi K2 (Jul 2025) leads

Claude 3.7 Sonnet: 18.6 (#277), Kimi K2 (Jul 2025): 23.3 (#179)

Reasoning benchmarks
BenchmarkClaude 3.7 SonnetKimi K2 (Jul 2025)
SimpleBench46.4%26.3%
LMArena Hard Prompts13331384
Epoch Capabilities Index141.16146.01
ForecastBench61.860.2
ARC-AGI-20.9%—
Kagi LLM Benchmark—64.4%
ARC-AGI-128.6%—
EnigmaEval4.2%—
LiveBench Reasoning87.8%—
LiveBench Data Analysis74%—
LiveBench76.1%—

Math Kimi K2 (Jul 2025) leads

Claude 3.7 Sonnet: 37.5 (#153), Kimi K2 (Jul 2025): 42.7 (#83)

Math benchmarks
BenchmarkClaude 3.7 SonnetKimi K2 (Jul 2025)
Omni-MATH33%65.4%
LMArena Math13371397
FrontierMath (Feb 2025 set)4.1%21.4%
OTIS Mock AIME 2024-202557.8%—
LiveBench Math79%—
MATH Level 591.2%—
FrontierMath Tier 4 (v1)—0%

Knowledge Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 39.8 (#130), Kimi K2 (Jul 2025): 37.3 (#157)

Knowledge benchmarks
BenchmarkClaude 3.7 SonnetKimi K2 (Jul 2025)
MMLU-Pro78.4%81.9%
Confabulations14.7%20.4%
GPQA (HELM)60.8%65.3%
LMArena Expert13211365
GPQA Diamond79.7%—
Humanity's Last Exam8%—
Vectara Hallucination Rate—17.9%

Multimodal Not comparable

Claude 3.7 Sonnet: 33.7 (#95), Kimi K2 (Jul 2025): —

Multimodal benchmarks
BenchmarkClaude 3.7 SonnetKimi K2 (Jul 2025)
LMArena Vision1169—
GeoBench68%—
VPCT39%—
SpatialViz-Bench33.9%—

Multilingual Kimi K2 (Jul 2025) leads

Claude 3.7 Sonnet: 44.1 (#179), Kimi K2 (Jul 2025): 49.6 (#130)

Multilingual benchmarks
BenchmarkClaude 3.7 SonnetKimi K2 (Jul 2025)
LMArena Non-English12961372
LMArena Chinese12991415
LMArena French13031379
LMArena German13011387
LMArena Japanese12671349
LMArena Korean12491325
LMArena Russian13111385
LMArena Spanish12981386

Instruction Following Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 72.9 (#125), Kimi K2 (Jul 2025): 71.1 (#156)

Instruction Following benchmarks
BenchmarkClaude 3.7 SonnetKimi K2 (Jul 2025)
IFEval83.4%85%
LMArena Instruction Following13521348
LiveBench Instruction Following81.3%—

Long Context Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 50.3 (#10), Kimi K2 (Jul 2025): 41.2 (#145)

Long Context benchmarks
BenchmarkClaude 3.7 SonnetKimi K2 (Jul 2025)
Fiction.LiveBench83.3%66.7%
LMArena Longer Query13731353
CL-bench—17.6%

Writing & Preference Kimi K2 (Jul 2025) leads

Claude 3.7 Sonnet: 54.4 (#150), Kimi K2 (Jul 2025): 62.3 (#78)

Writing & Preference benchmarks
BenchmarkClaude 3.7 SonnetKimi K2 (Jul 2025)
LMArena Text13141380
LMArena Creative Writing13321350
Short-Story Creative Writing81.1%85.6%
EQ-Bench Creative Writing14121666
WildBench81.4%86.2%
LMArena Multi-Turn13391371
LiveBench Language59.9%—

Frequently asked questions

Is Claude 3.7 Sonnet better than Kimi K2 (Jul 2025)?

Kimi K2 (Jul 2025) is the stronger model overall, scoring 41.2 to 39.5 on the Noometry Index.

Is Claude 3.7 Sonnet or Kimi K2 (Jul 2025) better for coding?

Kimi K2 (Jul 2025) scores higher on coding benchmarks: 42.4 versus 40.6 in the Noometry coding category.

How many benchmarks do Claude 3.7 Sonnet and Kimi K2 (Jul 2025) share?

34 benchmarks have published results for both models. Claude 3.7 Sonnet has 58 scored results on Noometry and Kimi K2 (Jul 2025) has 42.

Related comparisons

Go deeper