Model comparison

DeepSeek-V2.5 (Sep 2024) vs Granite 3.1 2b Instruct

DeepSeek-V2.5 (Sep 2024) is the stronger model overall, scoring 37.6 to 33.2 on the Noometry Index.

Last verified . 12 shared benchmarks.

DeepSeek-V2.5 (Sep 2024) DeepSeek

37.6

Rank #200 Confirmed

Granite 3.1 2b Instruct IBM

33.2

Rank #247 Confirmed

Summary

  • They share 12 benchmarks with published results for both. DeepSeek-V2.5 (Sep 2024) scores higher in 7 categories and Granite 3.1 2b Instruct in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where DeepSeek-V2.5 (Sep 2024) leads 49.8 to 34.1.

Side by side

DeepSeek-V2.5 (Sep 2024) and Granite 3.1 2b Instruct specifications
DeepSeek-V2.5 (Sep 2024)Granite 3.1 2b Instruct
ProviderDeepSeekIBM
Noometry Index37.633.2
Released2024-09-06—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2212

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Granite 3.1 2b Instruct leads

DeepSeek-V2.5 (Sep 2024): 31.7 (#281), Granite 3.1 2b Instruct: 33.4 (#257)

Coding benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Granite 3.1 2b Instruct
LMArena Coding13091149
Aider Polyglot17.8%—
BigCodeBench Instruct48.6%—
BigCodeBench Complete53.2%—
HumanEval+83.5%—
MBPP+74.1%—

Reasoning DeepSeek-V2.5 (Sep 2024) leads

DeepSeek-V2.5 (Sep 2024): 25.6 (#145), Granite 3.1 2b Instruct: 22.0 (#209)

Reasoning benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Granite 3.1 2b Instruct
LMArena Hard Prompts12891138

Math DeepSeek-V2.5 (Sep 2024) leads

DeepSeek-V2.5 (Sep 2024): 35.9 (#177), Granite 3.1 2b Instruct: 33.1 (#206)

Math benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Granite 3.1 2b Instruct
LMArena Math12881159

Knowledge DeepSeek-V2.5 (Sep 2024) leads

DeepSeek-V2.5 (Sep 2024): 34.8 (#193), Granite 3.1 2b Instruct: 30.8 (#224)

Knowledge benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Granite 3.1 2b Instruct
LMArena Expert12661131

Multilingual DeepSeek-V2.5 (Sep 2024) leads

DeepSeek-V2.5 (Sep 2024): 42.5 (#193), Granite 3.1 2b Instruct: 29.1 (#269)

Multilingual benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Granite 3.1 2b Instruct
LMArena Non-English12731068
LMArena Chinese13181139
LMArena Russian12891063
LMArena French1289—
LMArena German1258—
LMArena Japanese1228—
LMArena Korean1209—
LMArena Spanish1248—

Instruction Following DeepSeek-V2.5 (Sep 2024) leads

DeepSeek-V2.5 (Sep 2024): 67.5 (#194), Granite 3.1 2b Instruct: 57.7 (#264)

Instruction Following benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Granite 3.1 2b Instruct
LMArena Instruction Following12801116

Long Context DeepSeek-V2.5 (Sep 2024) leads

DeepSeek-V2.5 (Sep 2024): 39.5 (#174), Granite 3.1 2b Instruct: 35.0 (#244)

Long Context benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Granite 3.1 2b Instruct
LMArena Longer Query13011155

Writing & Preference DeepSeek-V2.5 (Sep 2024) leads

DeepSeek-V2.5 (Sep 2024): 49.8 (#187), Granite 3.1 2b Instruct: 34.1 (#274)

Writing & Preference benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Granite 3.1 2b Instruct
LMArena Text12941127
LMArena Creative Writing12851116
LMArena Multi-Turn12971099

Frequently asked questions

Is DeepSeek-V2.5 (Sep 2024) better than Granite 3.1 2b Instruct?

DeepSeek-V2.5 (Sep 2024) is the stronger model overall, scoring 37.6 to 33.2 on the Noometry Index.

Is DeepSeek-V2.5 (Sep 2024) or Granite 3.1 2b Instruct better for coding?

Granite 3.1 2b Instruct scores higher on coding benchmarks: 33.4 versus 31.7 in the Noometry coding category.

How many benchmarks do DeepSeek-V2.5 (Sep 2024) and Granite 3.1 2b Instruct share?

12 benchmarks have published results for both models. DeepSeek-V2.5 (Sep 2024) has 22 scored results on Noometry and Granite 3.1 2b Instruct has 12.

Related comparisons

Go deeper