Model comparison

Claude Sonnet 4.6 vs DeepSeek V4 Pro

DeepSeek V4 Pro is the stronger model overall, scoring 54.3 to 50.3 on the Noometry Index.

Last verified . 42 shared benchmarks.

Claude Sonnet 4.6 Anthropic

50.3

Rank #50 Confirmed

DeepSeek V4 Pro DeepSeek

54.3

Rank #31 Confirmed

Summary

  • They share 42 benchmarks with published results for both. Claude Sonnet 4.6 scores higher in 5 categories and DeepSeek V4 Pro in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where DeepSeek V4 Pro leads 64.8 to 52.9.
  • The biggest single-benchmark swing is Chess Puzzles: 13% for Claude Sonnet 4.6 and 47% for DeepSeek V4 Pro.
  • DeepSeek V4 Pro is cheaper at $0.66 / $1.98 per million input/output tokens, against $3 / $15 for Claude Sonnet 4.6.
  • DeepSeek V4 Pro has downloadable open weights; the other is API-only.

Side by side

Claude Sonnet 4.6 and DeepSeek V4 Pro specifications
Claude Sonnet 4.6DeepSeek V4 Pro
ProviderAnthropicDeepSeek
Noometry Index50.354.3
Released2026-02-172026-04-24
WeightsProprietaryOpen
Context window1M1M
Max output128K393K
Input $ / M tokens$3$0.66
Output $ / M tokens$15$1.98
Results tracked5748

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek V4 Pro leads

Claude Sonnet 4.6: 46.3 (#67), DeepSeek V4 Pro: 52.4 (#34)

Coding benchmarks
BenchmarkClaude Sonnet 4.6DeepSeek V4 Pro
SWE-bench Verified75.2%77.6%
FrontierCode24.3%28.6%
LMArena WebDev15221582
SciCode46.8%51%
WeirdML66.1%66.2%
LMArena Coding15041470
ALE-Bench1,3271,403
DeepSWE29.9%—

Agentic & Tool Use Claude Sonnet 4.6 leads

Claude Sonnet 4.6: 39.1 (#28), DeepSeek V4 Pro: 32.8 (#58)

Agentic & Tool Use benchmarks
BenchmarkClaude Sonnet 4.6DeepSeek V4 Pro
APEX-Agents43%47.3%
Vending-Bench 27,2043,285
Terminal-Bench53.4%—
OSWorld 2.09.3%—
DeepResearch Bench54.9%—
OSWorld72.1%—
ExploitBench23.6%—
GBAEval48.8%—
GDP.pdf18%—
LMArena Search1221—

Reasoning DeepSeek V4 Pro leads

Claude Sonnet 4.6: 46.1 (#45), DeepSeek V4 Pro: 56.5 (#24)

Reasoning benchmarks
BenchmarkClaude Sonnet 4.6DeepSeek V4 Pro
ARC-AGI-260.4%61.3%
NYT Connections (extended)80.9%91.3%
ARC-AGI-186.5%90.5%
CritPt3.1%18%
Chess Puzzles13%47%
LMArena Hard Prompts14841461
Mystery Game Puzzles16%43%
DTBench89.9%93.9%
LMCA46.5%45.5%
Epoch Capabilities Index152.24155.31
ForecastBench6256.1
Kagi LLM Benchmark—53.5%
Thematic Generalization76.3%—
Surface Evolver Bench—40%

Math DeepSeek V4 Pro leads

Claude Sonnet 4.6: 52.9 (#49), DeepSeek V4 Pro: 64.8 (#30)

Math benchmarks
BenchmarkClaude Sonnet 4.6DeepSeek V4 Pro
OTIS Mock AIME 2024-202585.8%98.6%
ProofBench45%50%
LMArena Math14621455
FrontierMath (Tiers 1-3)—64.6%
FrontierMath Tier 4—26.8%
MathArena Final-Answer Competitions—76.6%
FrontierMath (Feb 2025 set)32.4%—
FrontierMath Tier 4 (v1)8.3%—

Knowledge DeepSeek V4 Pro leads

Claude Sonnet 4.6: 51.7 (#65), DeepSeek V4 Pro: 59.5 (#31)

Knowledge benchmarks
BenchmarkClaude Sonnet 4.6DeepSeek V4 Pro
GPQA Diamond87.4%91.7%
SimpleQA Verified35.5%52.9%
Vectara Hallucination Rate10.6%8.6%
LMArena Expert15001464

Multimodal Not comparable

Claude Sonnet 4.6: 38.0 (#68), DeepSeek V4 Pro: —

Multimodal benchmarks
BenchmarkClaude Sonnet 4.6DeepSeek V4 Pro
LMArena Vision1283—
Blueprint-Bench 26.7%—
LMArena Document1482—

Multilingual Too close to call

Claude Sonnet 4.6: 54.4 (#41), DeepSeek V4 Pro: 54.4 (#45)

Multilingual benchmarks
BenchmarkClaude Sonnet 4.6DeepSeek V4 Pro
LMArena Non-English14401439
LMArena Chinese14911486
LMArena French14651472
LMArena German14281458
LMArena Japanese14201445
LMArena Korean14111447
LMArena Russian14401453
LMArena Spanish14641458

Instruction Following Claude Sonnet 4.6 leads

Claude Sonnet 4.6: 77.4 (#25), DeepSeek V4 Pro: 76.1 (#47)

Instruction Following benchmarks
BenchmarkClaude Sonnet 4.6DeepSeek V4 Pro
LMArena Instruction Following14751448

Long Context Too close to call

Claude Sonnet 4.6: 45.3 (#44), DeepSeek V4 Pro: 45.0 (#51)

Long Context benchmarks
BenchmarkClaude Sonnet 4.6DeepSeek V4 Pro
LMArena Longer Query14791458
CL-bench Life—13.5%

Writing & Preference Claude Sonnet 4.6 leads

Claude Sonnet 4.6: 70.2 (#22), DeepSeek V4 Pro: 65.5 (#46)

Writing & Preference benchmarks
BenchmarkClaude Sonnet 4.6DeepSeek V4 Pro
LMArena Text14581451
LMArena Creative Writing14351446
EQ-Bench Creative Writing18101553
EQ-Bench 412071166
LMArena Multi-Turn14641467

Frequently asked questions

Is Claude Sonnet 4.6 better than DeepSeek V4 Pro?

DeepSeek V4 Pro is the stronger model overall, scoring 54.3 to 50.3 on the Noometry Index.

Which is cheaper, Claude Sonnet 4.6 or DeepSeek V4 Pro?

DeepSeek V4 Pro is cheaper. It lists at $0.66 per million input tokens and $1.98 per million output tokens; Claude Sonnet 4.6 lists at $3 and $15.

Is Claude Sonnet 4.6 or DeepSeek V4 Pro better for coding?

DeepSeek V4 Pro scores higher on coding benchmarks: 52.4 versus 46.3 in the Noometry coding category.

Which has the bigger context window?

Both accept 1M tokens.

How many benchmarks do Claude Sonnet 4.6 and DeepSeek V4 Pro share?

42 benchmarks have published results for both models. Claude Sonnet 4.6 has 57 scored results on Noometry and DeepSeek V4 Pro has 48.

Related comparisons

Go deeper