Model comparison

Deepseek Coder v2 vs Llama 3.1 Nemotron 70b Instruct

Llama 3.1 Nemotron 70b Instruct is the stronger model overall, scoring 37.6 to 35.9 on the Noometry Index.

Last verified . 14 shared benchmarks.

Deepseek Coder v2 DeepSeek

35.9

Rank #220 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Deepseek Coder v2 scores higher in 1 category and Llama 3.1 Nemotron 70b Instruct in 7 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Llama 3.1 Nemotron 70b Instruct leads 48.4 to 38.2.
  • The biggest single-benchmark swing is BigCodeBench Complete: 59.7% for Deepseek Coder v2 and 48.2% for Llama 3.1 Nemotron 70b Instruct.

Side by side

Deepseek Coder v2 and Llama 3.1 Nemotron 70b Instruct specifications
Deepseek Coder v2Llama 3.1 Nemotron 70b Instruct
ProviderDeepSeekNVIDIA
Noometry Index35.937.6
Released2024-06-172024-12-18
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2414

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Deepseek Coder v2 leads

Deepseek Coder v2: 38.1 (#183), Llama 3.1 Nemotron 70b Instruct: 35.9 (#216)

Coding benchmarks
BenchmarkDeepseek Coder v2Llama 3.1 Nemotron 70b Instruct
BigCodeBench Instruct48.2%38.7%
LMArena Coding12511272
BigCodeBench Complete59.7%48.2%
HumanEval+82.3%—
MBPP+75.1%—

Reasoning Llama 3.1 Nemotron 70b Instruct leads

Deepseek Coder v2: 23.6 (#176), Llama 3.1 Nemotron 70b Instruct: 25.0 (#152)

Reasoning benchmarks
BenchmarkDeepseek Coder v2Llama 3.1 Nemotron 70b Instruct
LMArena Hard Prompts12071266
WinoGrande83.7%—

Math Too close to call

Deepseek Coder v2: 34.9 (#190), Llama 3.1 Nemotron 70b Instruct: 35.5 (#182)

Math benchmarks
BenchmarkDeepseek Coder v2Llama 3.1 Nemotron 70b Instruct
LMArena Math12411271
GSM8K94.5%—

Knowledge Llama 3.1 Nemotron 70b Instruct leads

Deepseek Coder v2: 32.3 (#212), Llama 3.1 Nemotron 70b Instruct: 34.1 (#199)

Knowledge benchmarks
BenchmarkDeepseek Coder v2Llama 3.1 Nemotron 70b Instruct
LMArena Expert11811242
ARC (AI2) Challenge64.3%—

Multilingual Llama 3.1 Nemotron 70b Instruct leads

Deepseek Coder v2: 36.3 (#240), Llama 3.1 Nemotron 70b Instruct: 40.5 (#217)

Multilingual benchmarks
BenchmarkDeepseek Coder v2Llama 3.1 Nemotron 70b Instruct
LMArena Non-English11821245
LMArena Chinese12011263
LMArena Russian11881227
LMArena French1185—
LMArena German1164—
LMArena Japanese1126—
LMArena Korean1104—
LMArena Spanish1153—

Instruction Following Llama 3.1 Nemotron 70b Instruct leads

Deepseek Coder v2: 61.7 (#242), Llama 3.1 Nemotron 70b Instruct: 65.9 (#213)

Instruction Following benchmarks
BenchmarkDeepseek Coder v2Llama 3.1 Nemotron 70b Instruct
LMArena Instruction Following11801252

Long Context Too close to call

Deepseek Coder v2: 37.0 (#224), Llama 3.1 Nemotron 70b Instruct: 37.6 (#215)

Long Context benchmarks
BenchmarkDeepseek Coder v2Llama 3.1 Nemotron 70b Instruct
LMArena Longer Query12191238

Writing & Preference Llama 3.1 Nemotron 70b Instruct leads

Deepseek Coder v2: 38.2 (#253), Llama 3.1 Nemotron 70b Instruct: 48.4 (#203)

Writing & Preference benchmarks
BenchmarkDeepseek Coder v2Llama 3.1 Nemotron 70b Instruct
LMArena Text11911283
LMArena Creative Writing11201269
LMArena Multi-Turn11771275

Frequently asked questions

Is Deepseek Coder v2 better than Llama 3.1 Nemotron 70b Instruct?

Llama 3.1 Nemotron 70b Instruct is the stronger model overall, scoring 37.6 to 35.9 on the Noometry Index.

Is Deepseek Coder v2 or Llama 3.1 Nemotron 70b Instruct better for coding?

Deepseek Coder v2 scores higher on coding benchmarks: 38.1 versus 35.9 in the Noometry coding category.

How many benchmarks do Deepseek Coder v2 and Llama 3.1 Nemotron 70b Instruct share?

14 benchmarks have published results for both models. Deepseek Coder v2 has 24 scored results on Noometry and Llama 3.1 Nemotron 70b Instruct has 14.

Related comparisons

Go deeper