Model comparison

DeepSeek-V2 (MoE-236B, May 2024) vs Mercury 2.5

Mercury 2.5 has enough public results to be ranked (#242); DeepSeek-V2 (MoE-236B, May 2024) does not yet, so treat this comparison as directional.

Last verified . 0 shared benchmarks.

Mercury 2.5 Inception

33.5

Rank #242 Reported

Summary

  • DeepSeek-V2 (MoE-236B, May 2024) has downloadable open weights; the other is API-only.

Side by side

DeepSeek-V2 (MoE-236B, May 2024) and Mercury 2.5 specifications
DeepSeek-V2 (MoE-236B, May 2024)Mercury 2.5
ProviderDeepSeekInception
Noometry Index40.333.5
Released2024-05-072026-09-08
WeightsOpenProprietary
Context window—260K
Max output—66K
Input $ / M tokens—$0.04
Output $ / M tokens—$0.15
Results tracked104

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

DeepSeek-V2 (MoE-236B, May 2024): 40.4 (#139), Mercury 2.5: 39.5 (#156)

Coding benchmarks
BenchmarkDeepSeek-V2 (MoE-236B, May 2024)Mercury 2.5
SciCode—38.5%
BigCodeBench Instruct48.9%—
BigCodeBench Complete59.4%—
ALE-Bench—301.65

Reasoning Not comparable

DeepSeek-V2 (MoE-236B, May 2024): —, Mercury 2.5: 22.4 (#193)

Reasoning benchmarks
BenchmarkDeepSeek-V2 (MoE-236B, May 2024)Mercury 2.5
CritPt—0%
BIG-Bench Hard78.8%—
Epoch Capabilities Index124.77—
HellaSwag87.1%—
PIQA83.9%—
WinoGrande86.3%—

Math Not comparable

DeepSeek-V2 (MoE-236B, May 2024): —, Mercury 2.5: 23.3 (#272)

Math benchmarks
BenchmarkDeepSeek-V2 (MoE-236B, May 2024)Mercury 2.5
ProofBench—3%

Knowledge Not comparable

DeepSeek-V2 (MoE-236B, May 2024): —, Mercury 2.5: —

Knowledge benchmarks
BenchmarkDeepSeek-V2 (MoE-236B, May 2024)Mercury 2.5
ARC (AI2) Challenge92.2%—
MMLU78.4%—
TriviaQA80%—

Frequently asked questions

Is DeepSeek-V2 (MoE-236B, May 2024) better than Mercury 2.5?

Mercury 2.5 has enough public results to be ranked (#242); DeepSeek-V2 (MoE-236B, May 2024) does not yet, so treat this comparison as directional.

Is DeepSeek-V2 (MoE-236B, May 2024) or Mercury 2.5 better for coding?

They score almost the same on coding (40.4 vs 39.5); test both on your own repository before choosing.

How many benchmarks do DeepSeek-V2 (MoE-236B, May 2024) and Mercury 2.5 share?

0 benchmarks have published results for both models. DeepSeek-V2 (MoE-236B, May 2024) has 10 scored results on Noometry and Mercury 2.5 has 4.

Related comparisons

Go deeper