Model comparison

DeepSeek LLM 67B vs Devstral Small 2505

Devstral Small 2505 is the stronger model overall, scoring 34.3 to 24.9 on the Noometry Index.

Last verified . 0 shared benchmarks.

DeepSeek LLM 67B DeepSeek

24.9

Rank #347 Confirmed

Devstral Small 2505 Mistral AI

34.3

Rank #233 Reported

Summary

  • The widest gap is in coding, where Devstral Small 2505 leads 38.9 to 31.9.

Side by side

DeepSeek LLM 67B and Devstral Small 2505 specifications
DeepSeek LLM 67BDevstral Small 2505
ProviderDeepSeekMistral AI
Noometry Index24.934.3
Released2023-11-292025-05-07
WeightsOpenOpen
Context window—128K
Max output—128K
Input $ / M tokens—$0.10
Output $ / M tokens—$0.30
Results tracked154

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Devstral Small 2505 leads

DeepSeek LLM 67B: 31.9 (#278), Devstral Small 2505: 38.9 (#166)

Coding benchmarks
BenchmarkDeepSeek LLM 67BDevstral Small 2505
SWE-bench Verified (bash only)—56.4%
SciCode—28.8%
LMArena Coding1096—

Reasoning Devstral Small 2505 leads

DeepSeek LLM 67B: 16.5 (#304), Devstral Small 2505: 19.7 (#252)

Reasoning benchmarks
BenchmarkDeepSeek LLM 67BDevstral Small 2505
Kagi LLM Benchmark—37.7%
CritPt—0%
Chess Puzzles0%—
LMArena Hard Prompts1070—
Epoch Capabilities Index110.5—

Math Not comparable

DeepSeek LLM 67B: 8.7 (#324), Devstral Small 2505: —

Math benchmarks
BenchmarkDeepSeek LLM 67BDevstral Small 2505
OTIS Mock AIME 2024-20250.8%—
LMArena Math1108—
MATH Level 56.4%—

Knowledge Not comparable

DeepSeek LLM 67B: 7.0 (#313), Devstral Small 2505: —

Knowledge benchmarks
BenchmarkDeepSeek LLM 67BDevstral Small 2505
GPQA Diamond24.6%—

Multilingual Not comparable

DeepSeek LLM 67B: 29.4 (#267), Devstral Small 2505: —

Multilingual benchmarks
BenchmarkDeepSeek LLM 67BDevstral Small 2505
LMArena Non-English1073—
LMArena Chinese1132—

Instruction Following Not comparable

DeepSeek LLM 67B: 55.4 (#277), Devstral Small 2505: —

Instruction Following benchmarks
BenchmarkDeepSeek LLM 67BDevstral Small 2505
LMArena Instruction Following1079—

Long Context Not comparable

DeepSeek LLM 67B: 33.1 (#265), Devstral Small 2505: —

Long Context benchmarks
BenchmarkDeepSeek LLM 67BDevstral Small 2505
LMArena Longer Query1092—

Writing & Preference Not comparable

DeepSeek LLM 67B: 31.6 (#282), Devstral Small 2505: —

Writing & Preference benchmarks
BenchmarkDeepSeek LLM 67BDevstral Small 2505
LMArena Text1105—
LMArena Creative Writing1067—
LMArena Multi-Turn1082—

Frequently asked questions

Is DeepSeek LLM 67B better than Devstral Small 2505?

Devstral Small 2505 is the stronger model overall, scoring 34.3 to 24.9 on the Noometry Index.

Is DeepSeek LLM 67B or Devstral Small 2505 better for coding?

Devstral Small 2505 scores higher on coding benchmarks: 38.9 versus 31.9 in the Noometry coding category.

How many benchmarks do DeepSeek LLM 67B and Devstral Small 2505 share?

0 benchmarks have published results for both models. DeepSeek LLM 67B has 15 scored results on Noometry and Devstral Small 2505 has 4.

Related comparisons

Go deeper