Model comparison

DeepSeek LLM 67B vs phi-3-medium 14B

phi-3-medium 14B is the stronger model overall, scoring 29.7 to 24.9 on the Noometry Index.

Last verified . 3 shared benchmarks.

DeepSeek LLM 67B DeepSeek

24.9

Rank #347 Confirmed

phi-3-medium 14B Microsoft

29.7

Rank #306 Reported

Summary

  • They share 3 benchmarks with published results for both. DeepSeek LLM 67B scores higher in 0 categories and phi-3-medium 14B in 3 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in math, where phi-3-medium 14B leads 27.3 to 8.7.
  • The biggest single-benchmark swing is MATH Level 5: 6.4% for DeepSeek LLM 67B and 17.6% for phi-3-medium 14B.

Side by side

DeepSeek LLM 67B and phi-3-medium 14B specifications
DeepSeek LLM 67Bphi-3-medium 14B
ProviderDeepSeekMicrosoft
Noometry Index24.929.7
Released2023-11-292024-04-23
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1513

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding phi-3-medium 14B leads

DeepSeek LLM 67B: 31.9 (#278), phi-3-medium 14B: 36.8 (#201)

Coding benchmarks
BenchmarkDeepSeek LLM 67Bphi-3-medium 14B
BigCodeBench Instruct—37.6%
LMArena Coding1096—
BigCodeBench Complete—48.7%

Reasoning Not comparable

DeepSeek LLM 67B: 16.5 (#304), phi-3-medium 14B: —

Reasoning benchmarks
BenchmarkDeepSeek LLM 67Bphi-3-medium 14B
Epoch Capabilities Index110.5121.23
Chess Puzzles0%—
LMArena Hard Prompts1070—
Adversarial NLI—55.8%
BIG-Bench Hard—81.4%
HellaSwag—82.4%
WinoGrande—81.5%

Math phi-3-medium 14B leads

DeepSeek LLM 67B: 8.7 (#324), phi-3-medium 14B: 27.3 (#250)

Math benchmarks
BenchmarkDeepSeek LLM 67Bphi-3-medium 14B
MATH Level 56.4%17.6%
OTIS Mock AIME 2024-20250.8%—
LMArena Math1108—

Knowledge phi-3-medium 14B leads

DeepSeek LLM 67B: 7.0 (#313), phi-3-medium 14B: 9.1 (#306)

Knowledge benchmarks
BenchmarkDeepSeek LLM 67Bphi-3-medium 14B
GPQA Diamond24.6%27.6%
ARC (AI2) Challenge—91.6%
MMLU—78%
OpenBookQA—87.4%
TriviaQA—73.9%

Multilingual Not comparable

DeepSeek LLM 67B: 29.4 (#267), phi-3-medium 14B: —

Multilingual benchmarks
BenchmarkDeepSeek LLM 67Bphi-3-medium 14B
LMArena Non-English1073—
LMArena Chinese1132—

Instruction Following Not comparable

DeepSeek LLM 67B: 55.4 (#277), phi-3-medium 14B: —

Instruction Following benchmarks
BenchmarkDeepSeek LLM 67Bphi-3-medium 14B
LMArena Instruction Following1079—

Long Context Not comparable

DeepSeek LLM 67B: 33.1 (#265), phi-3-medium 14B: —

Long Context benchmarks
BenchmarkDeepSeek LLM 67Bphi-3-medium 14B
LMArena Longer Query1092—

Writing & Preference Not comparable

DeepSeek LLM 67B: 31.6 (#282), phi-3-medium 14B: —

Writing & Preference benchmarks
BenchmarkDeepSeek LLM 67Bphi-3-medium 14B
LMArena Text1105—
LMArena Creative Writing1067—
LMArena Multi-Turn1082—

Frequently asked questions

Is DeepSeek LLM 67B better than phi-3-medium 14B?

phi-3-medium 14B is the stronger model overall, scoring 29.7 to 24.9 on the Noometry Index.

Is DeepSeek LLM 67B or phi-3-medium 14B better for coding?

phi-3-medium 14B scores higher on coding benchmarks: 36.8 versus 31.9 in the Noometry coding category.

How many benchmarks do DeepSeek LLM 67B and phi-3-medium 14B share?

3 benchmarks have published results for both models. DeepSeek LLM 67B has 15 scored results on Noometry and phi-3-medium 14B has 13.

Related comparisons

Go deeper