Model comparison

Claude Sonnet 4 vs DeepSeek-R1

DeepSeek-R1 is the stronger model overall, scoring 42.3 to 40.8 on the Noometry Index.

Last verified . 43 shared benchmarks.

Claude Sonnet 4 Anthropic

40.8

Rank #145 Confirmed

DeepSeek-R1 DeepSeek

42.3

Rank #115 Confirmed

Summary

  • They share 43 benchmarks with published results for both. Claude Sonnet 4 scores higher in 2 categories and DeepSeek-R1 in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in long context, where DeepSeek-R1 leads 45.4 to 33.7.
  • The biggest single-benchmark swing is Fiction.LiveBench: 46.9% for Claude Sonnet 4 and 75% for DeepSeek-R1.
  • DeepSeek-R1 is cheaper at $0.50 / $2.15 per million input/output tokens, against $3 / $15 for Claude Sonnet 4.
  • Claude Sonnet 4 accepts more context: 200K tokens versus 164K.

Side by side

Claude Sonnet 4 and DeepSeek-R1 specifications
Claude Sonnet 4DeepSeek-R1
ProviderAnthropicDeepSeek
Noometry Index40.842.3
Released2025-05-222025-01-20
WeightsProprietaryProprietary
Context window200K164K
Max output64K64K
Input $ / M tokens$3$0.50
Output $ / M tokens$15$2.15
Results tracked5852

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-R1 leads

Claude Sonnet 4: 43.5 (#88), DeepSeek-R1: 46.3 (#68)

Coding benchmarks
BenchmarkClaude Sonnet 4DeepSeek-R1
Aider Polyglot61.3%71.4%
SciCode40%35.7%
WeirdML46.1%41.6%
LMArena Coding14141427
ALE-Bench655.35804.12
SWE-bench Verified (bash only)64.9%—
GSO4.9%—
LiveBench Coding—66.7%
AlgoTune—1.7

Agentic & Tool Use Claude Sonnet 4 leads

Claude Sonnet 4: 38.5 (#31), DeepSeek-R1: 30.7 (#75)

Agentic & Tool Use benchmarks
BenchmarkClaude Sonnet 4DeepSeek-R1
DeepResearch Bench46.6%35.1%
METR Time Horizons62%53.8%
TheAgentCompany33.1%—
Cybench35%—
OSWorld43.9%—
BALROG—34.9%

Reasoning Claude Sonnet 4 leads

Claude Sonnet 4: 22.9 (#187), DeepSeek-R1: 18.6 (#278)

Reasoning benchmarks
BenchmarkClaude Sonnet 4DeepSeek-R1
ARC-AGI-25.9%1.3%
SimpleBench45.5%40.8%
Kagi LLM Benchmark73%69.4%
ARC-AGI-140%21.2%
CritPt0.3%1.1%
LMArena Hard Prompts13721416
Epoch Capabilities Index141.69141.29
ForecastBench60.260
EnigmaEval3.1%—
LiveBench Reasoning—83.2%
DTBench77.1%—
LiveBench Data Analysis—69.8%
LMCA29%—
LiveBench—71.6%

Math Too close to call

Claude Sonnet 4: 43.3 (#80), DeepSeek-R1: 43.8 (#79)

Math benchmarks
BenchmarkClaude Sonnet 4DeepSeek-R1
OTIS Mock AIME 2024-202571.1%66.4%
Omni-MATH60.2%42.4%
LMArena Math13751400
MATH Level 584.4%96.6%
LiveBench Math—80.7%
FrontierMath (Feb 2025 set)4.1%—
FrontierMath Tier 4 (v1)0%—

Knowledge DeepSeek-R1 leads

Claude Sonnet 4: 41.8 (#108), DeepSeek-R1: 44.5 (#87)

Knowledge benchmarks
BenchmarkClaude Sonnet 4DeepSeek-R1
GPQA Diamond79.2%76.3%
MMLU-Pro84.3%79.3%
Confabulations13.2%12.7%
Vectara Hallucination Rate10.3%11.3%
GPQA (HELM)70.6%66.6%
LMArena Expert13721394
Humanity's Last Exam7.8%—

Multimodal Not comparable

Claude Sonnet 4: 26.2 (#121), DeepSeek-R1: —

Multimodal benchmarks
BenchmarkClaude Sonnet 4DeepSeek-R1
LMArena Vision1191—
GeoBench37%—
VPCT34%—
MindCube44.8%—

Multilingual DeepSeek-R1 leads

Claude Sonnet 4: 46.7 (#156), DeepSeek-R1: 52.4 (#85)

Multilingual benchmarks
BenchmarkClaude Sonnet 4DeepSeek-R1
LMArena Non-English13331412
LMArena Chinese13501442
LMArena French13631417
LMArena German13311404
LMArena Japanese13021391
LMArena Korean12911360
LMArena Russian13551423
LMArena Spanish13571411

Instruction Following Too close to call

Claude Sonnet 4: 71.7 (#145), DeepSeek-R1: 72.0 (#143)

Instruction Following benchmarks
BenchmarkClaude Sonnet 4DeepSeek-R1
IFEval84%78.4%
LMArena Instruction Following13761382
LiveBench Instruction Following—80.5%

Long Context DeepSeek-R1 leads

Claude Sonnet 4: 33.7 (#259), DeepSeek-R1: 45.4 (#36)

Long Context benchmarks
BenchmarkClaude Sonnet 4DeepSeek-R1
Fiction.LiveBench46.9%75%
LMArena Longer Query13981391

Writing & Preference DeepSeek-R1 leads

Claude Sonnet 4: 57.1 (#132), DeepSeek-R1: 61.4 (#88)

Writing & Preference benchmarks
BenchmarkClaude Sonnet 4DeepSeek-R1
LMArena Text13511428
LMArena Creative Writing13451405
Short-Story Creative Writing81.4%83%
EQ-Bench Creative Writing14831500
WildBench83.8%82.8%
LMArena Multi-Turn13761405
LiveBench Language—48.5%

Frequently asked questions

Is Claude Sonnet 4 better than DeepSeek-R1?

DeepSeek-R1 is the stronger model overall, scoring 42.3 to 40.8 on the Noometry Index.

Which is cheaper, Claude Sonnet 4 or DeepSeek-R1?

DeepSeek-R1 is cheaper. It lists at $0.50 per million input tokens and $2.15 per million output tokens; Claude Sonnet 4 lists at $3 and $15.

Is Claude Sonnet 4 or DeepSeek-R1 better for coding?

DeepSeek-R1 scores higher on coding benchmarks: 46.3 versus 43.5 in the Noometry coding category.

Which has the bigger context window?

Claude Sonnet 4 does, with 200K tokens against 164K.

How many benchmarks do Claude Sonnet 4 and DeepSeek-R1 share?

43 benchmarks have published results for both models. Claude Sonnet 4 has 58 scored results on Noometry and DeepSeek-R1 has 52.

Related comparisons

Go deeper