Model comparison

DeepSeek-V2 (MoE-236B, May 2024) vs Phi-2

Neither DeepSeek-V2 (MoE-236B, May 2024) nor Phi-2 has enough public benchmark results to be ranked yet; the rows below show what has been published.

Last verified . 7 shared benchmarks.

Summary

  • They share 7 benchmarks with published results for both.

Side by side

DeepSeek-V2 (MoE-236B, May 2024) and Phi-2 specifications
DeepSeek-V2 (MoE-236B, May 2024)Phi-2
ProviderDeepSeekMicrosoft
Noometry Index40.3—
Released2024-05-072023-12-12
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1011

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

DeepSeek-V2 (MoE-236B, May 2024): 40.4 (#139), Phi-2: —

Coding benchmarks
BenchmarkDeepSeek-V2 (MoE-236B, May 2024)Phi-2
BigCodeBench Instruct48.9%—
BigCodeBench Complete59.4%—
HumanEval+—45.1%
MBPP+—54.2%

Reasoning Not comparable

DeepSeek-V2 (MoE-236B, May 2024): —, Phi-2: —

Reasoning benchmarks
BenchmarkDeepSeek-V2 (MoE-236B, May 2024)Phi-2
BIG-Bench Hard78.8%59.4%
Epoch Capabilities Index124.77107.94
HellaSwag87.1%53.6%
WinoGrande86.3%54.7%
Adversarial NLI—42.5%
PIQA83.9%—

Knowledge Not comparable

DeepSeek-V2 (MoE-236B, May 2024): —, Phi-2: —

Knowledge benchmarks
BenchmarkDeepSeek-V2 (MoE-236B, May 2024)Phi-2
ARC (AI2) Challenge92.2%75.9%
MMLU78.4%58.4%
TriviaQA80%45.2%
OpenBookQA—73.6%

Frequently asked questions

Is DeepSeek-V2 (MoE-236B, May 2024) better than Phi-2?

Neither DeepSeek-V2 (MoE-236B, May 2024) nor Phi-2 has enough public benchmark results to be ranked yet; the rows below show what has been published.

How many benchmarks do DeepSeek-V2 (MoE-236B, May 2024) and Phi-2 share?

7 benchmarks have published results for both models. DeepSeek-V2 (MoE-236B, May 2024) has 10 scored results on Noometry and Phi-2 has 11.

Related comparisons

Go deeper