Model comparison

Deepseek Coder v2 vs Magistral Small

Deepseek Coder v2 is the stronger model overall, scoring 35.9 to 30.2 on the Noometry Index.

Last verified . 0 shared benchmarks.

Deepseek Coder v2 DeepSeek

35.9

Rank #220 Confirmed

Magistral Small Mistral AI

30.2

Rank #296 Confirmed

Summary

  • The widest gap is in reasoning, where Deepseek Coder v2 leads 23.6 to 6.8.

Side by side

Deepseek Coder v2 and Magistral Small specifications
Deepseek Coder v2Magistral Small
ProviderDeepSeekMistral AI
Noometry Index35.930.2
Released2024-06-172025-06-10
WeightsOpenOpen
Context window—128K
Max output—40K
Input $ / M tokens—$0.50
Output $ / M tokens—$1.50
Results tracked2410

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Deepseek Coder v2: 38.1 (#183), Magistral Small: 38.4 (#176)

Coding benchmarks
BenchmarkDeepseek Coder v2Magistral Small
SciCode—35.2%
BigCodeBench Instruct48.2%—
LMArena Coding1251—
BigCodeBench Complete59.7%—
HumanEval+82.3%—
MBPP+75.1%—

Reasoning Deepseek Coder v2 leads

Deepseek Coder v2: 23.6 (#176), Magistral Small: 6.8 (#350)

Reasoning benchmarks
BenchmarkDeepseek Coder v2Magistral Small
ARC-AGI-2—0%
Kagi LLM Benchmark—6.3%
ARC-AGI-1—5%
CritPt—0.3%
Chess Puzzles—3%
LMArena Hard Prompts1207—
DTBench—61.3%
Epoch Capabilities Index—133.19
WinoGrande83.7%—

Math Deepseek Coder v2 leads

Deepseek Coder v2: 34.9 (#190), Magistral Small: 26.2 (#261)

Math benchmarks
BenchmarkDeepseek Coder v2Magistral Small
OTIS Mock AIME 2024-2025—30%
LMArena Math1241—
GSM8K94.5%—

Knowledge Deepseek Coder v2 leads

Deepseek Coder v2: 32.3 (#212), Magistral Small: 30.9 (#223)

Knowledge benchmarks
BenchmarkDeepseek Coder v2Magistral Small
GPQA Diamond—56.1%
LMArena Expert1181—
ARC (AI2) Challenge64.3%—

Multilingual Not comparable

Deepseek Coder v2: 36.3 (#240), Magistral Small: —

Multilingual benchmarks
BenchmarkDeepseek Coder v2Magistral Small
LMArena Non-English1182—
LMArena Chinese1201—
LMArena French1185—
LMArena German1164—
LMArena Japanese1126—
LMArena Korean1104—
LMArena Russian1188—
LMArena Spanish1153—

Instruction Following Not comparable

Deepseek Coder v2: 61.7 (#242), Magistral Small: —

Instruction Following benchmarks
BenchmarkDeepseek Coder v2Magistral Small
LMArena Instruction Following1180—

Long Context Not comparable

Deepseek Coder v2: 37.0 (#224), Magistral Small: —

Long Context benchmarks
BenchmarkDeepseek Coder v2Magistral Small
LMArena Longer Query1219—

Writing & Preference Not comparable

Deepseek Coder v2: 38.2 (#253), Magistral Small: —

Writing & Preference benchmarks
BenchmarkDeepseek Coder v2Magistral Small
LMArena Text1191—
LMArena Creative Writing1120—
LMArena Multi-Turn1177—

Frequently asked questions

Is Deepseek Coder v2 better than Magistral Small?

Deepseek Coder v2 is the stronger model overall, scoring 35.9 to 30.2 on the Noometry Index.

Is Deepseek Coder v2 or Magistral Small better for coding?

They score almost the same on coding (38.1 vs 38.4); test both on your own repository before choosing.

How many benchmarks do Deepseek Coder v2 and Magistral Small share?

0 benchmarks have published results for both models. Deepseek Coder v2 has 24 scored results on Noometry and Magistral Small has 10.

Related comparisons

Go deeper