Model comparison

Claude 3.5 Sonnet vs DeepSeek Coder 1.3B

Claude 3.5 Sonnet has enough public results to be ranked (#231); DeepSeek Coder 1.3B does not yet, so treat this comparison as directional.

Last verified . 6 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

DeepSeek Coder 1.3B DeepSeek

35.0

Unranked Sparse

Summary

  • They share 6 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 1 category and DeepSeek Coder 1.3B in 0 categories; one gap is clear of the uncertainty.
  • The widest gap is in coding, where Claude 3.5 Sonnet leads 39.0 to 31.2.
  • The biggest single-benchmark swing is BigCodeBench Complete: 58.6% for Claude 3.5 Sonnet and 29.6% for DeepSeek Coder 1.3B.
  • DeepSeek Coder 1.3B has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Sonnet and DeepSeek Coder 1.3B specifications
Claude 3.5 SonnetDeepSeek Coder 1.3B
ProviderAnthropicDeepSeek
Noometry Index34.635.0
Released2024-06-202023-11-02
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked609

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 39.0 (#165), DeepSeek Coder 1.3B: 31.2 (#287)

Coding benchmarks
BenchmarkClaude 3.5 SonnetDeepSeek Coder 1.3B
BigCodeBench Instruct46.8%22.8%
BigCodeBench Complete58.6%29.6%
HumanEval+81.7%60.4%
MBPP+74.3%54.8%
Aider Polyglot51.6%—
GSO4.6%—
WeirdML40%—
LiveBench Coding67.1%—
LMArena Coding1342—
CadEval48%—

Agentic & Tool Use Not comparable

Claude 3.5 Sonnet: 32.3 (#67), DeepSeek Coder 1.3B: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 SonnetDeepSeek Coder 1.3B
TheAgentCompany24%—
Cybench17.5%—
BALROG32.6%—
METR Time Horizons45.2%—

Reasoning Not comparable

Claude 3.5 Sonnet: 23.1 (#183), DeepSeek Coder 1.3B: —

Reasoning benchmarks
BenchmarkClaude 3.5 SonnetDeepSeek Coder 1.3B
Epoch Capabilities Index133.5563.6
SimpleBench41.4%—
EnigmaEval0.9%—
LiveBench Reasoning56.7%—
LMArena Hard Prompts1305—
DTBench67.8%—
LiveBench Data Analysis55%—
ForecastBench60.7—
LiveBench59%—
WinoGrande—53.3%

Math Not comparable

Claude 3.5 Sonnet: 19.2 (#288), DeepSeek Coder 1.3B: —

Math benchmarks
BenchmarkClaude 3.5 SonnetDeepSeek Coder 1.3B
OTIS Mock AIME 2024-20258.5%—
Omni-MATH27.6%—
LiveBench Math52.3%—
LMArena Math1307—
MATH Level 556.9%—
FrontierMath (Feb 2025 set)2.1%—
FrontierMath Tier 4 (v1)0%—
GSM8K—4.4%

Knowledge Not comparable

Claude 3.5 Sonnet: 28.6 (#245), DeepSeek Coder 1.3B: —

Knowledge benchmarks
BenchmarkClaude 3.5 SonnetDeepSeek Coder 1.3B
MMLU87.3%25.8%
GPQA Diamond55.3%—
Humanity's Last Exam4.1%—
MMLU-Pro77.7%—
Confabulations19.9%—
GPQA (HELM)56.5%—
LMArena Expert1265—
ARC (AI2) Challenge—25.4%

Multimodal Not comparable

Claude 3.5 Sonnet: 26.5 (#120), DeepSeek Coder 1.3B: —

Multimodal benchmarks
BenchmarkClaude 3.5 SonnetDeepSeek Coder 1.3B
LMArena Vision1125—
Video-MME60%—
GeoBench62%—
VPCT33%—

Multilingual Not comparable

Claude 3.5 Sonnet: 43.2 (#185), DeepSeek Coder 1.3B: —

Multilingual benchmarks
BenchmarkClaude 3.5 SonnetDeepSeek Coder 1.3B
LMArena Non-English1283—
LMArena Chinese1272—
LMArena French1305—
LMArena German1297—
LMArena Japanese1234—
LMArena Korean1200—
LMArena Russian1306—
LMArena Spanish1290—

Instruction Following Not comparable

Claude 3.5 Sonnet: 68.8 (#182), DeepSeek Coder 1.3B: —

Instruction Following benchmarks
BenchmarkClaude 3.5 SonnetDeepSeek Coder 1.3B
LiveBench Instruction Following69.3%—
IFEval85.5%—
LMArena Instruction Following1297—

Long Context Not comparable

Claude 3.5 Sonnet: 39.9 (#167), DeepSeek Coder 1.3B: —

Long Context benchmarks
BenchmarkClaude 3.5 SonnetDeepSeek Coder 1.3B
LMArena Longer Query1311—

Writing & Preference Not comparable

Claude 3.5 Sonnet: 52.9 (#164), DeepSeek Coder 1.3B: —

Writing & Preference benchmarks
BenchmarkClaude 3.5 SonnetDeepSeek Coder 1.3B
LMArena Text1298—
LMArena Creative Writing1292—
Short-Story Creative Writing80.3%—
EQ-Bench Creative Writing1451—
WildBench79.2%—
LMArena Multi-Turn1326—
LiveBench Language53.8%—

Frequently asked questions

Is Claude 3.5 Sonnet better than DeepSeek Coder 1.3B?

Claude 3.5 Sonnet has enough public results to be ranked (#231); DeepSeek Coder 1.3B does not yet, so treat this comparison as directional.

Is Claude 3.5 Sonnet or DeepSeek Coder 1.3B better for coding?

Claude 3.5 Sonnet scores higher on coding benchmarks: 39.0 versus 31.2 in the Noometry coding category.

How many benchmarks do Claude 3.5 Sonnet and DeepSeek Coder 1.3B share?

6 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and DeepSeek Coder 1.3B has 9.

Related comparisons

Go deeper