Model comparison

Codellama 70b Instruct vs DeepSeek-V2 (MoE-236B, May 2024)

Codellama 70b Instruct has enough public results to be ranked (#237); DeepSeek-V2 (MoE-236B, May 2024) does not yet, so treat this comparison as directional.

Last verified . 2 shared benchmarks.

Codellama 70b Instruct Meta

33.7

Rank #237 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Codellama 70b Instruct scores higher in 0 categories and DeepSeek-V2 (MoE-236B, May 2024) in 1 category; one gap is clear of the uncertainty.
  • The biggest single-benchmark swing is BigCodeBench Complete: 49.6% for Codellama 70b Instruct and 59.4% for DeepSeek-V2 (MoE-236B, May 2024).

Side by side

Codellama 70b Instruct and DeepSeek-V2 (MoE-236B, May 2024) specifications
Codellama 70b InstructDeepSeek-V2 (MoE-236B, May 2024)
ProviderMetaDeepSeek
Noometry Index33.740.3
Released—2024-05-07
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked710

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V2 (MoE-236B, May 2024) leads

Codellama 70b Instruct: 37.6 (#193), DeepSeek-V2 (MoE-236B, May 2024): 40.4 (#139)

Coding benchmarks
BenchmarkCodellama 70b InstructDeepSeek-V2 (MoE-236B, May 2024)
BigCodeBench Instruct40.7%48.9%
BigCodeBench Complete49.6%59.4%
HumanEval+65.9%—

Reasoning Not comparable

Codellama 70b Instruct: 20.1 (#242), DeepSeek-V2 (MoE-236B, May 2024): —

Reasoning benchmarks
BenchmarkCodellama 70b InstructDeepSeek-V2 (MoE-236B, May 2024)
LMArena Hard Prompts1052—
BIG-Bench Hard—78.8%
Epoch Capabilities Index—124.77
HellaSwag—87.1%
PIQA—83.9%
WinoGrande—86.3%

Knowledge Not comparable

Codellama 70b Instruct: —, DeepSeek-V2 (MoE-236B, May 2024): —

Knowledge benchmarks
BenchmarkCodellama 70b InstructDeepSeek-V2 (MoE-236B, May 2024)
ARC (AI2) Challenge—92.2%
MMLU—78.4%
TriviaQA—80%

Multilingual Not comparable

Codellama 70b Instruct: 24.8 (#288), DeepSeek-V2 (MoE-236B, May 2024): —

Multilingual benchmarks
BenchmarkCodellama 70b InstructDeepSeek-V2 (MoE-236B, May 2024)
LMArena Non-English992—

Instruction Following Not comparable

Codellama 70b Instruct: 51.9 (#293), DeepSeek-V2 (MoE-236B, May 2024): —

Instruction Following benchmarks
BenchmarkCodellama 70b InstructDeepSeek-V2 (MoE-236B, May 2024)
LMArena Instruction Following1024—

Writing & Preference Not comparable

Codellama 70b Instruct: 33.4 (#277), DeepSeek-V2 (MoE-236B, May 2024): —

Writing & Preference benchmarks
BenchmarkCodellama 70b InstructDeepSeek-V2 (MoE-236B, May 2024)
LMArena Text1057—

Frequently asked questions

Is Codellama 70b Instruct better than DeepSeek-V2 (MoE-236B, May 2024)?

Codellama 70b Instruct has enough public results to be ranked (#237); DeepSeek-V2 (MoE-236B, May 2024) does not yet, so treat this comparison as directional.

Is Codellama 70b Instruct or DeepSeek-V2 (MoE-236B, May 2024) better for coding?

DeepSeek-V2 (MoE-236B, May 2024) scores higher on coding benchmarks: 40.4 versus 37.6 in the Noometry coding category.

How many benchmarks do Codellama 70b Instruct and DeepSeek-V2 (MoE-236B, May 2024) share?

2 benchmarks have published results for both models. Codellama 70b Instruct has 7 scored results on Noometry and DeepSeek-V2 (MoE-236B, May 2024) has 10.

Related comparisons

Go deeper