Model comparison

DeepSeek-V2.5 (Sep 2024) vs Llama 2-70B

DeepSeek-V2.5 (Sep 2024) is the stronger model overall, scoring 37.6 to 24.4 on the Noometry Index.

Last verified . 17 shared benchmarks.

DeepSeek-V2.5 (Sep 2024) DeepSeek

37.6

Rank #200 Confirmed

Llama 2-70B Meta

24.4

Rank #349 Confirmed

Summary

  • They share 17 benchmarks with published results for both. DeepSeek-V2.5 (Sep 2024) scores higher in 8 categories and Llama 2-70B in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where DeepSeek-V2.5 (Sep 2024) leads 35.9 to 8.1.

Side by side

DeepSeek-V2.5 (Sep 2024) and Llama 2-70B specifications
DeepSeek-V2.5 (Sep 2024)Llama 2-70B
ProviderDeepSeekMeta
Noometry Index37.624.4
Released2024-09-062023-07-18
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2235

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

DeepSeek-V2.5 (Sep 2024): 31.7 (#281), Llama 2-70B: 31.4 (#286)

Coding benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Llama 2-70B
LMArena Coding13091079
Aider Polyglot17.8%—
BigCodeBench Instruct48.6%—
BigCodeBench Complete53.2%—
HumanEval+83.5%—
MBPP+74.1%—

Reasoning DeepSeek-V2.5 (Sep 2024) leads

DeepSeek-V2.5 (Sep 2024): 25.6 (#145), Llama 2-70B: 14.4 (#325)

Reasoning benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Llama 2-70B
LMArena Hard Prompts12891073
DTBench—41.6%
BIG-Bench Hard—64.9%
CommonsenseQA 2.0—50%
Epoch Capabilities Index—113.79
ForecastBench—51.4
HellaSwag—85.3%
LAMBADA—78.9%
PIQA—82.8%
WinoGrande—80.2%

Math DeepSeek-V2.5 (Sep 2024) leads

DeepSeek-V2.5 (Sep 2024): 35.9 (#177), Llama 2-70B: 8.1 (#326)

Math benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Llama 2-70B
LMArena Math12881091
OTIS Mock AIME 2024-2025—0%
MATH Level 5—3.3%
GSM8K—69.6%

Knowledge DeepSeek-V2.5 (Sep 2024) leads

DeepSeek-V2.5 (Sep 2024): 34.8 (#193), Llama 2-70B: 7.4 (#310)

Knowledge benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Llama 2-70B
LMArena Expert12661039
GPQA Diamond—26.3%
ARC (AI2) Challenge—78.3%
BoolQ—88.6%
MMLU—69.9%
OpenBookQA—60.2%
TriviaQA—87.6%

Multilingual DeepSeek-V2.5 (Sep 2024) leads

DeepSeek-V2.5 (Sep 2024): 42.5 (#193), Llama 2-70B: 27.7 (#274)

Multilingual benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Llama 2-70B
LMArena Non-English12731045
LMArena Chinese1318995
LMArena French12891090
LMArena German12581041
LMArena Japanese1228927
LMArena Korean1209964
LMArena Russian12891083
LMArena Spanish12481143

Instruction Following DeepSeek-V2.5 (Sep 2024) leads

DeepSeek-V2.5 (Sep 2024): 67.5 (#194), Llama 2-70B: 54.9 (#278)

Instruction Following benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Llama 2-70B
LMArena Instruction Following12801071

Long Context DeepSeek-V2.5 (Sep 2024) leads

DeepSeek-V2.5 (Sep 2024): 39.5 (#174), Llama 2-70B: 32.3 (#270)

Long Context benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Llama 2-70B
LMArena Longer Query13011062

Writing & Preference DeepSeek-V2.5 (Sep 2024) leads

DeepSeek-V2.5 (Sep 2024): 49.8 (#187), Llama 2-70B: 32.3 (#279)

Writing & Preference benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Llama 2-70B
LMArena Text12941115
LMArena Creative Writing12851075
LMArena Multi-Turn12971088

Frequently asked questions

Is DeepSeek-V2.5 (Sep 2024) better than Llama 2-70B?

DeepSeek-V2.5 (Sep 2024) is the stronger model overall, scoring 37.6 to 24.4 on the Noometry Index.

Is DeepSeek-V2.5 (Sep 2024) or Llama 2-70B better for coding?

They score almost the same on coding (31.7 vs 31.4); test both on your own repository before choosing.

How many benchmarks do DeepSeek-V2.5 (Sep 2024) and Llama 2-70B share?

17 benchmarks have published results for both models. DeepSeek-V2.5 (Sep 2024) has 22 scored results on Noometry and Llama 2-70B has 35.

Related comparisons

Go deeper