Model comparison

Claude 3.7 Sonnet vs DeepSeek-R1

DeepSeek-R1 is the stronger model overall, scoring 42.3 to 39.5 on the Noometry Index.

Last verified . 44 shared benchmarks.

Claude 3.7 Sonnet Anthropic

39.5

Rank #164 Confirmed

DeepSeek-R1 DeepSeek

42.3

Rank #115 Confirmed

Summary

  • They share 44 benchmarks with published results for both. Claude 3.7 Sonnet scores higher in 3 categories and DeepSeek-R1 in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where DeepSeek-R1 leads 52.4 to 44.1.
  • The biggest single-benchmark swing is LiveBench Language: 59.9% for Claude 3.7 Sonnet and 48.5% for DeepSeek-R1.

Side by side

Claude 3.7 Sonnet and DeepSeek-R1 specifications
Claude 3.7 SonnetDeepSeek-R1
ProviderAnthropicDeepSeek
Noometry Index39.542.3
Released2025-02-242025-01-20
WeightsProprietaryProprietary
Context window—164K
Max output—64K
Input $ / M tokens—$0.50
Output $ / M tokens—$2.15
Results tracked5852

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-R1 leads

Claude 3.7 Sonnet: 40.6 (#136), DeepSeek-R1: 46.3 (#68)

Coding benchmarks
BenchmarkClaude 3.7 SonnetDeepSeek-R1
Aider Polyglot64.9%71.4%
LiveBench Coding74.5%66.7%
LMArena Coding13611427
SWE-bench Verified61%—
SWE-bench Verified (bash only)52.8%—
SciCode—35.7%
GSO3.8%—
WeirdML—41.6%
CadEval54%—
ALE-Bench—804.12
AlgoTune—1.7

Agentic & Tool Use Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 34.1 (#50), DeepSeek-R1: 30.7 (#75)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.7 SonnetDeepSeek-R1
DeepResearch Bench43.6%35.1%
METR Time Horizons60%53.8%
TheAgentCompany30.9%—
Cybench20%—
OSWorld35.8%—
BALROG—34.9%

Reasoning Too close to call

Claude 3.7 Sonnet: 18.6 (#277), DeepSeek-R1: 18.6 (#278)

Reasoning benchmarks
BenchmarkClaude 3.7 SonnetDeepSeek-R1
ARC-AGI-20.9%1.3%
SimpleBench46.4%40.8%
ARC-AGI-128.6%21.2%
LiveBench Reasoning87.8%83.2%
LMArena Hard Prompts13331416
LiveBench Data Analysis74%69.8%
Epoch Capabilities Index141.16141.29
ForecastBench61.860
LiveBench76.1%71.6%
Kagi LLM Benchmark—69.4%
CritPt—1.1%
EnigmaEval4.2%—

Math DeepSeek-R1 leads

Claude 3.7 Sonnet: 37.5 (#153), DeepSeek-R1: 43.8 (#79)

Math benchmarks
BenchmarkClaude 3.7 SonnetDeepSeek-R1
OTIS Mock AIME 2024-202557.8%66.4%
Omni-MATH33%42.4%
LiveBench Math79%80.7%
LMArena Math13371400
MATH Level 591.2%96.6%
FrontierMath (Feb 2025 set)4.1%—

Knowledge DeepSeek-R1 leads

Claude 3.7 Sonnet: 39.8 (#130), DeepSeek-R1: 44.5 (#87)

Knowledge benchmarks
BenchmarkClaude 3.7 SonnetDeepSeek-R1
GPQA Diamond79.7%76.3%
MMLU-Pro78.4%79.3%
Confabulations14.7%12.7%
GPQA (HELM)60.8%66.6%
LMArena Expert13211394
Humanity's Last Exam8%—
Vectara Hallucination Rate—11.3%

Multimodal Not comparable

Claude 3.7 Sonnet: 33.7 (#95), DeepSeek-R1: —

Multimodal benchmarks
BenchmarkClaude 3.7 SonnetDeepSeek-R1
LMArena Vision1169—
GeoBench68%—
VPCT39%—
SpatialViz-Bench33.9%—

Multilingual DeepSeek-R1 leads

Claude 3.7 Sonnet: 44.1 (#179), DeepSeek-R1: 52.4 (#85)

Multilingual benchmarks
BenchmarkClaude 3.7 SonnetDeepSeek-R1
LMArena Non-English12961412
LMArena Chinese12991442
LMArena French13031417
LMArena German13011404
LMArena Japanese12671391
LMArena Korean12491360
LMArena Russian13111423
LMArena Spanish12981411

Instruction Following Too close to call

Claude 3.7 Sonnet: 72.9 (#125), DeepSeek-R1: 72.0 (#143)

Instruction Following benchmarks
BenchmarkClaude 3.7 SonnetDeepSeek-R1
LiveBench Instruction Following81.3%80.5%
IFEval83.4%78.4%
LMArena Instruction Following13521382

Long Context Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 50.3 (#10), DeepSeek-R1: 45.4 (#36)

Long Context benchmarks
BenchmarkClaude 3.7 SonnetDeepSeek-R1
Fiction.LiveBench83.3%75%
LMArena Longer Query13731391

Writing & Preference DeepSeek-R1 leads

Claude 3.7 Sonnet: 54.4 (#150), DeepSeek-R1: 61.4 (#88)

Writing & Preference benchmarks
BenchmarkClaude 3.7 SonnetDeepSeek-R1
LMArena Text13141428
LMArena Creative Writing13321405
Short-Story Creative Writing81.1%83%
EQ-Bench Creative Writing14121500
WildBench81.4%82.8%
LMArena Multi-Turn13391405
LiveBench Language59.9%48.5%

Frequently asked questions

Is Claude 3.7 Sonnet better than DeepSeek-R1?

DeepSeek-R1 is the stronger model overall, scoring 42.3 to 39.5 on the Noometry Index.

Is Claude 3.7 Sonnet or DeepSeek-R1 better for coding?

DeepSeek-R1 scores higher on coding benchmarks: 46.3 versus 40.6 in the Noometry coding category.

How many benchmarks do Claude 3.7 Sonnet and DeepSeek-R1 share?

44 benchmarks have published results for both models. Claude 3.7 Sonnet has 58 scored results on Noometry and DeepSeek-R1 has 52.

Related comparisons

Go deeper