Model comparison

DeepSeek-V2.5 (Sep 2024) vs Granite 4.1 8b

DeepSeek-V2.5 (Sep 2024) and Granite 4.1 8b score almost the same on the Noometry Index (37.6 vs 37.4), so choose on price, context window or the category you care about most.

Last verified . 12 shared benchmarks.

DeepSeek-V2.5 (Sep 2024) DeepSeek

37.6

Rank #200 Confirmed

Granite 4.1 8b IBM

37.4

Rank #205 Confirmed

Summary

  • They share 12 benchmarks with published results for both. DeepSeek-V2.5 (Sep 2024) scores higher in 5 categories and Granite 4.1 8b in 3 categories; 3 gaps are clear of the uncertainty.

Side by side

DeepSeek-V2.5 (Sep 2024) and Granite 4.1 8b specifications
DeepSeek-V2.5 (Sep 2024)Granite 4.1 8b
ProviderDeepSeekIBM
Noometry Index37.637.4
Released2024-09-06—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2213

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V2.5 (Sep 2024) leads

DeepSeek-V2.5 (Sep 2024): 31.7 (#281), Granite 4.1 8b: 30.1 (#297)

Coding benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Granite 4.1 8b
LMArena Coding13091312
Aider Polyglot17.8%—
LMArena WebDev—1192
BigCodeBench Instruct48.6%—
BigCodeBench Complete53.2%—
HumanEval+83.5%—
MBPP+74.1%—

Reasoning Too close to call

DeepSeek-V2.5 (Sep 2024): 25.6 (#145), Granite 4.1 8b: 25.7 (#143)

Reasoning benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Granite 4.1 8b
LMArena Hard Prompts12891293

Math Too close to call

DeepSeek-V2.5 (Sep 2024): 35.9 (#177), Granite 4.1 8b: 36.4 (#166)

Math benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Granite 4.1 8b
LMArena Math12881312

Knowledge Granite 4.1 8b leads

DeepSeek-V2.5 (Sep 2024): 34.8 (#193), Granite 4.1 8b: 36.1 (#174)

Knowledge benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Granite 4.1 8b
LMArena Expert12661309

Multilingual Too close to call

DeepSeek-V2.5 (Sep 2024): 42.5 (#193), Granite 4.1 8b: 41.7 (#204)

Multilingual benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Granite 4.1 8b
LMArena Non-English12731261
LMArena Chinese13181337
LMArena Russian12891240
LMArena French1289—
LMArena German1258—
LMArena Japanese1228—
LMArena Korean1209—
LMArena Spanish1248—

Instruction Following Too close to call

DeepSeek-V2.5 (Sep 2024): 67.5 (#194), Granite 4.1 8b: 66.9 (#203)

Instruction Following benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Granite 4.1 8b
LMArena Instruction Following12801269

Long Context Too close to call

DeepSeek-V2.5 (Sep 2024): 39.5 (#174), Granite 4.1 8b: 38.7 (#193)

Long Context benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Granite 4.1 8b
LMArena Longer Query13011275

Writing & Preference DeepSeek-V2.5 (Sep 2024) leads

DeepSeek-V2.5 (Sep 2024): 49.8 (#187), Granite 4.1 8b: 48.1 (#204)

Writing & Preference benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Granite 4.1 8b
LMArena Text12941290
LMArena Creative Writing12851251
LMArena Multi-Turn12971268

Frequently asked questions

Is DeepSeek-V2.5 (Sep 2024) better than Granite 4.1 8b?

DeepSeek-V2.5 (Sep 2024) and Granite 4.1 8b score almost the same on the Noometry Index (37.6 vs 37.4), so choose on price, context window or the category you care about most.

Is DeepSeek-V2.5 (Sep 2024) or Granite 4.1 8b better for coding?

DeepSeek-V2.5 (Sep 2024) scores higher on coding benchmarks: 31.7 versus 30.1 in the Noometry coding category.

How many benchmarks do DeepSeek-V2.5 (Sep 2024) and Granite 4.1 8b share?

12 benchmarks have published results for both models. DeepSeek-V2.5 (Sep 2024) has 22 scored results on Noometry and Granite 4.1 8b has 13.

Related comparisons

Go deeper