Model comparison

Claude 3.5 Sonnet vs DeepSeek Coder 33B

Claude 3.5 Sonnet has enough public results to be ranked (#231); DeepSeek Coder 33B does not yet, so treat this comparison as directional.

Last verified . 6 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

DeepSeek Coder 33B DeepSeek

38.9

Unranked Sparse

Summary

  • They share 6 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 1 category and DeepSeek Coder 33B in 0 categories, but none of those gaps is larger than the uncertainty.
  • The biggest single-benchmark swing is BigCodeBench Complete: 58.6% for Claude 3.5 Sonnet and 51.1% for DeepSeek Coder 33B.
  • DeepSeek Coder 33B has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Sonnet and DeepSeek Coder 33B specifications
Claude 3.5 SonnetDeepSeek Coder 33B
ProviderAnthropicDeepSeek
Noometry Index34.638.9
Released2024-06-202023-11-02
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked609

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Claude 3.5 Sonnet: 39.0 (#165), DeepSeek Coder 33B: 38.0 (#184)

Coding benchmarks
BenchmarkClaude 3.5 SonnetDeepSeek Coder 33B
BigCodeBench Instruct46.8%42%
BigCodeBench Complete58.6%51.1%
HumanEval+81.7%75%
MBPP+74.3%70.1%
Aider Polyglot51.6%—
GSO4.6%—
WeirdML40%—
LiveBench Coding67.1%—
LMArena Coding1342—
CadEval48%—

Agentic & Tool Use Not comparable

Claude 3.5 Sonnet: 32.3 (#67), DeepSeek Coder 33B: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 SonnetDeepSeek Coder 33B
TheAgentCompany24%—
Cybench17.5%—
BALROG32.6%—
METR Time Horizons45.2%—

Reasoning Not comparable

Claude 3.5 Sonnet: 23.1 (#183), DeepSeek Coder 33B: —

Reasoning benchmarks
BenchmarkClaude 3.5 SonnetDeepSeek Coder 33B
Epoch Capabilities Index133.5596.32
SimpleBench41.4%—
EnigmaEval0.9%—
LiveBench Reasoning56.7%—
LMArena Hard Prompts1305—
DTBench67.8%—
LiveBench Data Analysis55%—
ForecastBench60.7—
LiveBench59%—
WinoGrande—62%

Math Not comparable

Claude 3.5 Sonnet: 19.2 (#288), DeepSeek Coder 33B: —

Math benchmarks
BenchmarkClaude 3.5 SonnetDeepSeek Coder 33B
OTIS Mock AIME 2024-20258.5%—
Omni-MATH27.6%—
LiveBench Math52.3%—
LMArena Math1307—
MATH Level 556.9%—
FrontierMath (Feb 2025 set)2.1%—
FrontierMath Tier 4 (v1)0%—
GSM8K—35.4%

Knowledge Not comparable

Claude 3.5 Sonnet: 28.6 (#245), DeepSeek Coder 33B: —

Knowledge benchmarks
BenchmarkClaude 3.5 SonnetDeepSeek Coder 33B
MMLU87.3%39.4%
GPQA Diamond55.3%—
Humanity's Last Exam4.1%—
MMLU-Pro77.7%—
Confabulations19.9%—
GPQA (HELM)56.5%—
LMArena Expert1265—
ARC (AI2) Challenge—42.2%

Multimodal Not comparable

Claude 3.5 Sonnet: 26.5 (#120), DeepSeek Coder 33B: —

Multimodal benchmarks
BenchmarkClaude 3.5 SonnetDeepSeek Coder 33B
LMArena Vision1125—
Video-MME60%—
GeoBench62%—
VPCT33%—

Multilingual Not comparable

Claude 3.5 Sonnet: 43.2 (#185), DeepSeek Coder 33B: —

Multilingual benchmarks
BenchmarkClaude 3.5 SonnetDeepSeek Coder 33B
LMArena Non-English1283—
LMArena Chinese1272—
LMArena French1305—
LMArena German1297—
LMArena Japanese1234—
LMArena Korean1200—
LMArena Russian1306—
LMArena Spanish1290—

Instruction Following Not comparable

Claude 3.5 Sonnet: 68.8 (#182), DeepSeek Coder 33B: —

Instruction Following benchmarks
BenchmarkClaude 3.5 SonnetDeepSeek Coder 33B
LiveBench Instruction Following69.3%—
IFEval85.5%—
LMArena Instruction Following1297—

Long Context Not comparable

Claude 3.5 Sonnet: 39.9 (#167), DeepSeek Coder 33B: —

Long Context benchmarks
BenchmarkClaude 3.5 SonnetDeepSeek Coder 33B
LMArena Longer Query1311—

Writing & Preference Not comparable

Claude 3.5 Sonnet: 52.9 (#164), DeepSeek Coder 33B: —

Writing & Preference benchmarks
BenchmarkClaude 3.5 SonnetDeepSeek Coder 33B
LMArena Text1298—
LMArena Creative Writing1292—
Short-Story Creative Writing80.3%—
EQ-Bench Creative Writing1451—
WildBench79.2%—
LMArena Multi-Turn1326—
LiveBench Language53.8%—

Frequently asked questions

Is Claude 3.5 Sonnet better than DeepSeek Coder 33B?

Claude 3.5 Sonnet has enough public results to be ranked (#231); DeepSeek Coder 33B does not yet, so treat this comparison as directional.

Is Claude 3.5 Sonnet or DeepSeek Coder 33B better for coding?

They score almost the same on coding (39.0 vs 38.0); test both on your own repository before choosing.

How many benchmarks do Claude 3.5 Sonnet and DeepSeek Coder 33B share?

6 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and DeepSeek Coder 33B has 9.

Related comparisons

Go deeper