Model comparison

Deepseek Coder v2 vs Llama 3.1 Tulu 3 8b

Deepseek Coder v2 and Llama 3.1 Tulu 3 8b score almost the same on the Noometry Index (35.9 vs 35.7), so choose on price, context window or the category you care about most.

Last verified . 11 shared benchmarks.

Deepseek Coder v2 DeepSeek

35.9

Rank #220 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Deepseek Coder v2 scores higher in 6 categories and Llama 3.1 Tulu 3 8b in 1 category; 4 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Deepseek Coder v2 leads 38.1 to 34.4.

Side by side

Deepseek Coder v2 and Llama 3.1 Tulu 3 8b specifications
Deepseek Coder v2Llama 3.1 Tulu 3 8b
ProviderDeepSeekAllen Institute for AI (Ai2)
Noometry Index35.935.7
Released2024-06-17—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2411

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Deepseek Coder v2 leads

Deepseek Coder v2: 38.1 (#183), Llama 3.1 Tulu 3 8b: 34.4 (#235)

Coding benchmarks
BenchmarkDeepseek Coder v2Llama 3.1 Tulu 3 8b
LMArena Coding12511183
BigCodeBench Instruct48.2%—
BigCodeBench Complete59.7%—
HumanEval+82.3%—
MBPP+75.1%—

Reasoning Too close to call

Deepseek Coder v2: 23.6 (#176), Llama 3.1 Tulu 3 8b: 22.8 (#188)

Reasoning benchmarks
BenchmarkDeepseek Coder v2Llama 3.1 Tulu 3 8b
LMArena Hard Prompts12071174
WinoGrande83.7%—

Math Deepseek Coder v2 leads

Deepseek Coder v2: 34.9 (#190), Llama 3.1 Tulu 3 8b: 33.9 (#198)

Math benchmarks
BenchmarkDeepseek Coder v2Llama 3.1 Tulu 3 8b
LMArena Math12411195
GSM8K94.5%—

Knowledge Not comparable

Deepseek Coder v2: 32.3 (#212), Llama 3.1 Tulu 3 8b: —

Knowledge benchmarks
BenchmarkDeepseek Coder v2Llama 3.1 Tulu 3 8b
LMArena Expert1181—
ARC (AI2) Challenge64.3%—

Multilingual Too close to call

Deepseek Coder v2: 36.3 (#240), Llama 3.1 Tulu 3 8b: 35.4 (#246)

Multilingual benchmarks
BenchmarkDeepseek Coder v2Llama 3.1 Tulu 3 8b
LMArena Non-English11821169
LMArena Chinese12011176
LMArena Russian11881193
LMArena French1185—
LMArena German1164—
LMArena Japanese1126—
LMArena Korean1104—
LMArena Spanish1153—

Instruction Following Too close to call

Deepseek Coder v2: 61.7 (#242), Llama 3.1 Tulu 3 8b: 61.3 (#246)

Instruction Following benchmarks
BenchmarkDeepseek Coder v2Llama 3.1 Tulu 3 8b
LMArena Instruction Following11801174

Long Context Deepseek Coder v2 leads

Deepseek Coder v2: 37.0 (#224), Llama 3.1 Tulu 3 8b: 35.8 (#239)

Long Context benchmarks
BenchmarkDeepseek Coder v2Llama 3.1 Tulu 3 8b
LMArena Longer Query12191181

Writing & Preference Llama 3.1 Tulu 3 8b leads

Deepseek Coder v2: 38.2 (#253), Llama 3.1 Tulu 3 8b: 39.7 (#245)

Writing & Preference benchmarks
BenchmarkDeepseek Coder v2Llama 3.1 Tulu 3 8b
LMArena Text11911193
LMArena Creative Writing11201182
LMArena Multi-Turn11771154

Frequently asked questions

Is Deepseek Coder v2 better than Llama 3.1 Tulu 3 8b?

Deepseek Coder v2 and Llama 3.1 Tulu 3 8b score almost the same on the Noometry Index (35.9 vs 35.7), so choose on price, context window or the category you care about most.

Is Deepseek Coder v2 or Llama 3.1 Tulu 3 8b better for coding?

Deepseek Coder v2 scores higher on coding benchmarks: 38.1 versus 34.4 in the Noometry coding category.

How many benchmarks do Deepseek Coder v2 and Llama 3.1 Tulu 3 8b share?

11 benchmarks have published results for both models. Deepseek Coder v2 has 24 scored results on Noometry and Llama 3.1 Tulu 3 8b has 11.

Related comparisons

Go deeper