Model comparison

DeepSeek-V2.5 (Sep 2024) vs Olmo 3.1 32b Instruct

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 37.6 on the Noometry Index.

Last verified . 16 shared benchmarks.

Summary

  • They share 16 benchmarks with published results for both. DeepSeek-V2.5 (Sep 2024) scores higher in 0 categories and Olmo 3.1 32b Instruct in 8 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Olmo 3.1 32b Instruct leads 39.5 to 31.7.

Side by side

DeepSeek-V2.5 (Sep 2024) and Olmo 3.1 32b Instruct specifications
DeepSeek-V2.5 (Sep 2024)Olmo 3.1 32b Instruct
ProviderDeepSeekAllen Institute for AI (Ai2)
Noometry Index37.639.4
Released2024-09-06—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2216

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Instruct leads

DeepSeek-V2.5 (Sep 2024): 31.7 (#281), Olmo 3.1 32b Instruct: 39.5 (#157)

Coding benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Olmo 3.1 32b Instruct
LMArena Coding13091347
Aider Polyglot17.8%—
BigCodeBench Instruct48.6%—
BigCodeBench Complete53.2%—
HumanEval+83.5%—
MBPP+74.1%—

Reasoning Too close to call

DeepSeek-V2.5 (Sep 2024): 25.6 (#145), Olmo 3.1 32b Instruct: 26.4 (#132)

Reasoning benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Olmo 3.1 32b Instruct
LMArena Hard Prompts12891322

Math Too close to call

DeepSeek-V2.5 (Sep 2024): 35.9 (#177), Olmo 3.1 32b Instruct: 36.3 (#167)

Math benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Olmo 3.1 32b Instruct
LMArena Math12881305

Knowledge Olmo 3.1 32b Instruct leads

DeepSeek-V2.5 (Sep 2024): 34.8 (#193), Olmo 3.1 32b Instruct: 36.1 (#175)

Knowledge benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Olmo 3.1 32b Instruct
LMArena Expert12661308

Multilingual Too close to call

DeepSeek-V2.5 (Sep 2024): 42.5 (#193), Olmo 3.1 32b Instruct: 42.6 (#191)

Multilingual benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Olmo 3.1 32b Instruct
LMArena Non-English12731275
LMArena Chinese13181304
LMArena French12891328
LMArena German12581282
LMArena Korean12091206
LMArena Russian12891268
LMArena Spanish12481336
LMArena Japanese1228—

Instruction Following Olmo 3.1 32b Instruct leads

DeepSeek-V2.5 (Sep 2024): 67.5 (#194), Olmo 3.1 32b Instruct: 68.6 (#187)

Instruction Following benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Olmo 3.1 32b Instruct
LMArena Instruction Following12801299

Long Context Too close to call

DeepSeek-V2.5 (Sep 2024): 39.5 (#174), Olmo 3.1 32b Instruct: 39.9 (#166)

Long Context benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Olmo 3.1 32b Instruct
LMArena Longer Query13011312

Writing & Preference Too close to call

DeepSeek-V2.5 (Sep 2024): 49.8 (#187), Olmo 3.1 32b Instruct: 50.2 (#185)

Writing & Preference benchmarks
BenchmarkDeepSeek-V2.5 (Sep 2024)Olmo 3.1 32b Instruct
LMArena Text12941311
LMArena Creative Writing12851264
LMArena Multi-Turn12971309

Frequently asked questions

Is DeepSeek-V2.5 (Sep 2024) better than Olmo 3.1 32b Instruct?

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 37.6 on the Noometry Index.

Is DeepSeek-V2.5 (Sep 2024) or Olmo 3.1 32b Instruct better for coding?

Olmo 3.1 32b Instruct scores higher on coding benchmarks: 39.5 versus 31.7 in the Noometry coding category.

How many benchmarks do DeepSeek-V2.5 (Sep 2024) and Olmo 3.1 32b Instruct share?

16 benchmarks have published results for both models. DeepSeek-V2.5 (Sep 2024) has 22 scored results on Noometry and Olmo 3.1 32b Instruct has 16.

Related comparisons

Go deeper