Model comparison

DeepSeek Coder 1.3B vs Llama2 70b Steerlm Chat

Llama2 70b Steerlm Chat has enough public results to be ranked (#268); DeepSeek Coder 1.3B does not yet, so treat this comparison as directional.

Last verified . 0 shared benchmarks.

DeepSeek Coder 1.3B DeepSeek

35.0

Unranked Sparse

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Side by side

DeepSeek Coder 1.3B and Llama2 70b Steerlm Chat specifications
DeepSeek Coder 1.3BLlama2 70b Steerlm Chat
ProviderDeepSeekNVIDIA
Noometry Index35.031.8
Released2023-11-02—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked99

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek Coder 1.3B leads

DeepSeek Coder 1.3B: 31.2 (#287), Llama2 70b Steerlm Chat: 29.9 (#300)

Coding benchmarks
BenchmarkDeepSeek Coder 1.3BLlama2 70b Steerlm Chat
BigCodeBench Instruct22.8%—
LMArena Coding—1025
BigCodeBench Complete29.6%—
HumanEval+60.4%—
MBPP+54.8%—

Reasoning Not comparable

DeepSeek Coder 1.3B: —, Llama2 70b Steerlm Chat: 20.0 (#246)

Reasoning benchmarks
BenchmarkDeepSeek Coder 1.3BLlama2 70b Steerlm Chat
LMArena Hard Prompts—1047
Epoch Capabilities Index63.6—
WinoGrande53.3%—

Math Not comparable

DeepSeek Coder 1.3B: —, Llama2 70b Steerlm Chat: 31.3 (#226)

Math benchmarks
BenchmarkDeepSeek Coder 1.3BLlama2 70b Steerlm Chat
LMArena Math—1072
GSM8K4.4%—

Knowledge Not comparable

DeepSeek Coder 1.3B: —, Llama2 70b Steerlm Chat: —

Knowledge benchmarks
BenchmarkDeepSeek Coder 1.3BLlama2 70b Steerlm Chat
ARC (AI2) Challenge25.4%—
MMLU25.8%—

Multilingual Not comparable

DeepSeek Coder 1.3B: —, Llama2 70b Steerlm Chat: 28.8 (#270)

Multilingual benchmarks
BenchmarkDeepSeek Coder 1.3BLlama2 70b Steerlm Chat
LMArena Non-English—1063

Instruction Following Not comparable

DeepSeek Coder 1.3B: —, Llama2 70b Steerlm Chat: 54.2 (#279)

Instruction Following benchmarks
BenchmarkDeepSeek Coder 1.3BLlama2 70b Steerlm Chat
LMArena Instruction Following—1060

Long Context Not comparable

DeepSeek Coder 1.3B: —, Llama2 70b Steerlm Chat: 30.4 (#288)

Long Context benchmarks
BenchmarkDeepSeek Coder 1.3BLlama2 70b Steerlm Chat
LMArena Longer Query—998

Writing & Preference Not comparable

DeepSeek Coder 1.3B: —, Llama2 70b Steerlm Chat: 31.6 (#283)

Writing & Preference benchmarks
BenchmarkDeepSeek Coder 1.3BLlama2 70b Steerlm Chat
LMArena Text—1098
LMArena Creative Writing—1091
LMArena Multi-Turn—1058

Frequently asked questions

Is DeepSeek Coder 1.3B better than Llama2 70b Steerlm Chat?

Llama2 70b Steerlm Chat has enough public results to be ranked (#268); DeepSeek Coder 1.3B does not yet, so treat this comparison as directional.

Is DeepSeek Coder 1.3B or Llama2 70b Steerlm Chat better for coding?

DeepSeek Coder 1.3B scores higher on coding benchmarks: 31.2 versus 29.9 in the Noometry coding category.

How many benchmarks do DeepSeek Coder 1.3B and Llama2 70b Steerlm Chat share?

0 benchmarks have published results for both models. DeepSeek Coder 1.3B has 9 scored results on Noometry and Llama2 70b Steerlm Chat has 9.

Related comparisons

Go deeper