Model comparison

Deepseek Coder v2 vs DeepSeek LLM 67B

Deepseek Coder v2 is the stronger model overall, scoring 35.9 to 24.9 on the Noometry Index.

Last verified . 10 shared benchmarks.

Deepseek Coder v2 DeepSeek

35.9

Rank #220 Confirmed

DeepSeek LLM 67B DeepSeek

24.9

Rank #347 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Deepseek Coder v2 scores higher in 8 categories and DeepSeek LLM 67B in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Deepseek Coder v2 leads 34.9 to 8.7.

Side by side

Deepseek Coder v2 and DeepSeek LLM 67B specifications
Deepseek Coder v2DeepSeek LLM 67B
ProviderDeepSeekDeepSeek
Noometry Index35.924.9
Released2024-06-172023-11-29
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2415

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Deepseek Coder v2 leads

Deepseek Coder v2: 38.1 (#183), DeepSeek LLM 67B: 31.9 (#278)

Coding benchmarks
BenchmarkDeepseek Coder v2DeepSeek LLM 67B
LMArena Coding12511096
BigCodeBench Instruct48.2%—
BigCodeBench Complete59.7%—
HumanEval+82.3%—
MBPP+75.1%—

Reasoning Deepseek Coder v2 leads

Deepseek Coder v2: 23.6 (#176), DeepSeek LLM 67B: 16.5 (#304)

Reasoning benchmarks
BenchmarkDeepseek Coder v2DeepSeek LLM 67B
LMArena Hard Prompts12071070
Chess Puzzles—0%
Epoch Capabilities Index—110.5
WinoGrande83.7%—

Math Deepseek Coder v2 leads

Deepseek Coder v2: 34.9 (#190), DeepSeek LLM 67B: 8.7 (#324)

Math benchmarks
BenchmarkDeepseek Coder v2DeepSeek LLM 67B
LMArena Math12411108
OTIS Mock AIME 2024-2025—0.8%
MATH Level 5—6.4%
GSM8K94.5%—

Knowledge Deepseek Coder v2 leads

Deepseek Coder v2: 32.3 (#212), DeepSeek LLM 67B: 7.0 (#313)

Knowledge benchmarks
BenchmarkDeepseek Coder v2DeepSeek LLM 67B
GPQA Diamond—24.6%
LMArena Expert1181—
ARC (AI2) Challenge64.3%—

Multilingual Deepseek Coder v2 leads

Deepseek Coder v2: 36.3 (#240), DeepSeek LLM 67B: 29.4 (#267)

Multilingual benchmarks
BenchmarkDeepseek Coder v2DeepSeek LLM 67B
LMArena Non-English11821073
LMArena Chinese12011132
LMArena French1185—
LMArena German1164—
LMArena Japanese1126—
LMArena Korean1104—
LMArena Russian1188—
LMArena Spanish1153—

Instruction Following Deepseek Coder v2 leads

Deepseek Coder v2: 61.7 (#242), DeepSeek LLM 67B: 55.4 (#277)

Instruction Following benchmarks
BenchmarkDeepseek Coder v2DeepSeek LLM 67B
LMArena Instruction Following11801079

Long Context Deepseek Coder v2 leads

Deepseek Coder v2: 37.0 (#224), DeepSeek LLM 67B: 33.1 (#265)

Long Context benchmarks
BenchmarkDeepseek Coder v2DeepSeek LLM 67B
LMArena Longer Query12191092

Writing & Preference Deepseek Coder v2 leads

Deepseek Coder v2: 38.2 (#253), DeepSeek LLM 67B: 31.6 (#282)

Writing & Preference benchmarks
BenchmarkDeepseek Coder v2DeepSeek LLM 67B
LMArena Text11911105
LMArena Creative Writing11201067
LMArena Multi-Turn11771082

Frequently asked questions

Is Deepseek Coder v2 better than DeepSeek LLM 67B?

Deepseek Coder v2 is the stronger model overall, scoring 35.9 to 24.9 on the Noometry Index.

Is Deepseek Coder v2 or DeepSeek LLM 67B better for coding?

Deepseek Coder v2 scores higher on coding benchmarks: 38.1 versus 31.9 in the Noometry coding category.

How many benchmarks do Deepseek Coder v2 and DeepSeek LLM 67B share?

10 benchmarks have published results for both models. Deepseek Coder v2 has 24 scored results on Noometry and DeepSeek LLM 67B has 15.

Related comparisons

Go deeper