Model comparison

Codellama 70b Instruct vs DeepSeek-V3.1-Terminus

DeepSeek-V3.1-Terminus is the stronger model overall, scoring 43.1 to 33.7 on the Noometry Index.

Last verified . 4 shared benchmarks.

Codellama 70b Instruct Meta

33.7

Rank #237 Confirmed

DeepSeek-V3.1-Terminus DeepSeek

43.1

Rank #97 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Codellama 70b Instruct scores higher in 0 categories and DeepSeek-V3.1-Terminus in 5 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where DeepSeek-V3.1-Terminus leads 61.0 to 33.4.

Side by side

Codellama 70b Instruct and DeepSeek-V3.1-Terminus specifications
Codellama 70b InstructDeepSeek-V3.1-Terminus
ProviderMetaDeepSeek
Noometry Index33.743.1
Released—2025-09-22
WeightsOpenOpen
Context window—164K
Max output—147K
Input $ / M tokens—$0.27
Output $ / M tokens—$1
Results tracked716

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3.1-Terminus leads

Codellama 70b Instruct: 37.6 (#193), DeepSeek-V3.1-Terminus: 42.0 (#113)

Coding benchmarks
BenchmarkCodellama 70b InstructDeepSeek-V3.1-Terminus
SciCode—40.6%
BigCodeBench Instruct40.7%—
LMArena Coding—1426
BigCodeBench Complete49.6%—
ALE-Bench—745.17
HumanEval+65.9%—

Reasoning DeepSeek-V3.1-Terminus leads

Codellama 70b Instruct: 20.1 (#242), DeepSeek-V3.1-Terminus: 26.4 (#133)

Reasoning benchmarks
BenchmarkCodellama 70b InstructDeepSeek-V3.1-Terminus
LMArena Hard Prompts10521426
Kagi LLM Benchmark—57.4%
CritPt—1.7%
DTBench—81.3%
LMCA—28.6%

Math Not comparable

Codellama 70b Instruct: —, DeepSeek-V3.1-Terminus: 38.5 (#137)

Math benchmarks
BenchmarkCodellama 70b InstructDeepSeek-V3.1-Terminus
LMArena Math—1402

Multilingual DeepSeek-V3.1-Terminus leads

Codellama 70b Instruct: 24.8 (#288), DeepSeek-V3.1-Terminus: 52.1 (#92)

Multilingual benchmarks
BenchmarkCodellama 70b InstructDeepSeek-V3.1-Terminus
LMArena Non-English9921407
LMArena Russian—1436

Instruction Following DeepSeek-V3.1-Terminus leads

Codellama 70b Instruct: 51.9 (#293), DeepSeek-V3.1-Terminus: 74.0 (#106)

Instruction Following benchmarks
BenchmarkCodellama 70b InstructDeepSeek-V3.1-Terminus
LMArena Instruction Following10241404

Long Context Not comparable

Codellama 70b Instruct: —, DeepSeek-V3.1-Terminus: 43.4 (#97)

Long Context benchmarks
BenchmarkCodellama 70b InstructDeepSeek-V3.1-Terminus
LMArena Longer Query—1421

Writing & Preference DeepSeek-V3.1-Terminus leads

Codellama 70b Instruct: 33.4 (#277), DeepSeek-V3.1-Terminus: 61.0 (#92)

Writing & Preference benchmarks
BenchmarkCodellama 70b InstructDeepSeek-V3.1-Terminus
LMArena Text10571419
LMArena Creative Writing—1403
LMArena Multi-Turn—1411

Frequently asked questions

Is Codellama 70b Instruct better than DeepSeek-V3.1-Terminus?

DeepSeek-V3.1-Terminus is the stronger model overall, scoring 43.1 to 33.7 on the Noometry Index.

Is Codellama 70b Instruct or DeepSeek-V3.1-Terminus better for coding?

DeepSeek-V3.1-Terminus scores higher on coding benchmarks: 42.0 versus 37.6 in the Noometry coding category.

How many benchmarks do Codellama 70b Instruct and DeepSeek-V3.1-Terminus share?

4 benchmarks have published results for both models. Codellama 70b Instruct has 7 scored results on Noometry and DeepSeek-V3.1-Terminus has 16.

Related comparisons

Go deeper