Model comparison

Codellama 70b Instruct vs DeepSeek-V3.1

DeepSeek-V3.1 is the stronger model overall, scoring 42.8 to 33.7 on the Noometry Index.

Last verified . 4 shared benchmarks.

Codellama 70b Instruct Meta

33.7

Rank #237 Confirmed

DeepSeek-V3.1 DeepSeek

42.8

Rank #108 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Codellama 70b Instruct scores higher in 0 categories and DeepSeek-V3.1 in 5 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where DeepSeek-V3.1 leads 60.3 to 33.4.

Side by side

Codellama 70b Instruct and DeepSeek-V3.1 specifications
Codellama 70b InstructDeepSeek-V3.1
ProviderMetaDeepSeek
Noometry Index33.742.8
Released—2025-08-21
WeightsOpenOpen
Context window—164K
Max output—8K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.95
Results tracked727

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3.1 leads

Codellama 70b Instruct: 37.6 (#193), DeepSeek-V3.1: 40.3 (#144)

Coding benchmarks
BenchmarkCodellama 70b InstructDeepSeek-V3.1
WeirdML—38.4%
BigCodeBench Instruct40.7%—
LMArena Coding—1417
BigCodeBench Complete49.6%—
HumanEval+65.9%—

Reasoning DeepSeek-V3.1 leads

Codellama 70b Instruct: 20.1 (#242), DeepSeek-V3.1: 27.9 (#110)

Reasoning benchmarks
BenchmarkCodellama 70b InstructDeepSeek-V3.1
LMArena Hard Prompts10521417
SimpleBench—40%
Kagi LLM Benchmark—53.2%
DTBench—82.7%
LMCA—24.3%
Epoch Capabilities Index—139.92
ForecastBench—58

Math Not comparable

Codellama 70b Instruct: —, DeepSeek-V3.1: 38.9 (#122)

Math benchmarks
BenchmarkCodellama 70b InstructDeepSeek-V3.1
LMArena Math—1420

Knowledge Not comparable

Codellama 70b Instruct: —, DeepSeek-V3.1: 43.7 (#90)

Knowledge benchmarks
BenchmarkCodellama 70b InstructDeepSeek-V3.1
Vectara Hallucination Rate—5.5%
LMArena Expert—1405

Multilingual DeepSeek-V3.1 leads

Codellama 70b Instruct: 24.8 (#288), DeepSeek-V3.1: 51.6 (#106)

Multilingual benchmarks
BenchmarkCodellama 70b InstructDeepSeek-V3.1
LMArena Non-English9921400
LMArena Chinese—1469
LMArena French—1447
LMArena German—1411
LMArena Japanese—1378
LMArena Korean—1337
LMArena Russian—1405
LMArena Spanish—1431

Instruction Following DeepSeek-V3.1 leads

Codellama 70b Instruct: 51.9 (#293), DeepSeek-V3.1: 73.9 (#110)

Instruction Following benchmarks
BenchmarkCodellama 70b InstructDeepSeek-V3.1
LMArena Instruction Following10241400

Long Context Not comparable

Codellama 70b Instruct: —, DeepSeek-V3.1: 36.3 (#232)

Long Context benchmarks
BenchmarkCodellama 70b InstructDeepSeek-V3.1
Fiction.LiveBench—52.8%
LMArena Longer Query—1422

Writing & Preference DeepSeek-V3.1 leads

Codellama 70b Instruct: 33.4 (#277), DeepSeek-V3.1: 60.3 (#98)

Writing & Preference benchmarks
BenchmarkCodellama 70b InstructDeepSeek-V3.1
LMArena Text10571420
LMArena Creative Writing—1401
EQ-Bench Creative Writing—1436
LMArena Multi-Turn—1408

Frequently asked questions

Is Codellama 70b Instruct better than DeepSeek-V3.1?

DeepSeek-V3.1 is the stronger model overall, scoring 42.8 to 33.7 on the Noometry Index.

Is Codellama 70b Instruct or DeepSeek-V3.1 better for coding?

DeepSeek-V3.1 scores higher on coding benchmarks: 40.3 versus 37.6 in the Noometry coding category.

How many benchmarks do Codellama 70b Instruct and DeepSeek-V3.1 share?

4 benchmarks have published results for both models. Codellama 70b Instruct has 7 scored results on Noometry and DeepSeek-V3.1 has 27.

Related comparisons

Go deeper