Model comparison

DeepSeek Coder 33B vs Mercury

Mercury has enough public results to be ranked (#199); DeepSeek Coder 33B does not yet, so treat this comparison as directional.

Last verified . 0 shared benchmarks.

DeepSeek Coder 33B DeepSeek

38.9

Unranked Sparse

Mercury Inception

37.6

Rank #199 Confirmed

Summary

  • DeepSeek Coder 33B has downloadable open weights; the other is API-only.

Side by side

DeepSeek Coder 33B and Mercury specifications
DeepSeek Coder 33BMercury
ProviderDeepSeekInception
Noometry Index38.937.6
Released2023-11-02—
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked99

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

DeepSeek Coder 33B: 38.0 (#184), Mercury: 38.7 (#170)

Coding benchmarks
BenchmarkDeepSeek Coder 33BMercury
BigCodeBench Instruct42%—
LMArena Coding—1322
BigCodeBench Complete51.1%—
HumanEval+75%—
MBPP+70.1%—

Reasoning Not comparable

DeepSeek Coder 33B: —, Mercury: 17.5 (#293)

Reasoning benchmarks
BenchmarkDeepSeek Coder 33BMercury
Kagi LLM Benchmark—21.6%
LMArena Hard Prompts—1285
Epoch Capabilities Index96.32—
WinoGrande62%—

Math Not comparable

DeepSeek Coder 33B: —, Mercury: —

Math benchmarks
BenchmarkDeepSeek Coder 33BMercury
GSM8K35.4%—

Knowledge Not comparable

DeepSeek Coder 33B: —, Mercury: —

Knowledge benchmarks
BenchmarkDeepSeek Coder 33BMercury
ARC (AI2) Challenge42.2%—
MMLU39.4%—

Multilingual Not comparable

DeepSeek Coder 33B: —, Mercury: 41.6 (#206)

Multilingual benchmarks
BenchmarkDeepSeek Coder 33BMercury
LMArena Non-English—1260

Instruction Following Not comparable

DeepSeek Coder 33B: —, Mercury: 65.2 (#224)

Instruction Following benchmarks
BenchmarkDeepSeek Coder 33BMercury
LMArena Instruction Following—1239

Long Context Not comparable

DeepSeek Coder 33B: —, Mercury: 38.4 (#198)

Long Context benchmarks
BenchmarkDeepSeek Coder 33BMercury
LMArena Longer Query—1266

Writing & Preference Not comparable

DeepSeek Coder 33B: —, Mercury: 46.2 (#221)

Writing & Preference benchmarks
BenchmarkDeepSeek Coder 33BMercury
LMArena Text—1282
LMArena Creative Writing—1191
LMArena Multi-Turn—1282

Frequently asked questions

Is DeepSeek Coder 33B better than Mercury?

Mercury has enough public results to be ranked (#199); DeepSeek Coder 33B does not yet, so treat this comparison as directional.

Is DeepSeek Coder 33B or Mercury better for coding?

They score almost the same on coding (38.0 vs 38.7); test both on your own repository before choosing.

How many benchmarks do DeepSeek Coder 33B and Mercury share?

0 benchmarks have published results for both models. DeepSeek Coder 33B has 9 scored results on Noometry and Mercury has 9.

Related comparisons

Go deeper