Model comparison

Codestral vs Magistral Medium

Magistral Medium is the stronger model overall, scoring 35.2 to 30.6 on the Noometry Index. Codestral costs 6.1× less per token, which makes it the better buy when Magistral Medium's lead doesn't matter for your workload.

Last verified . 1 shared benchmarks.

Codestral Mistral AI

30.6

Rank #290 Reported

Magistral Medium Mistral AI

35.2

Rank #227 Confirmed

Summary

  • They share 1 benchmark with published results for both. Codestral scores higher in 1 category and Magistral Medium in 1 category; 2 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Magistral Medium leads 39.1 to 27.3.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 32.5% for Codestral and 16.2% for Magistral Medium.
  • Codestral is cheaper at $0.30 / $0.90 per million input/output tokens, against $2 / $5 for Magistral Medium.
  • Magistral Medium accepts more context: 262K tokens versus 256K.
  • Magistral Medium has downloadable open weights; the other is API-only.

Side by side

Codestral and Magistral Medium specifications
CodestralMagistral Medium
ProviderMistral AIMistral AI
Noometry Index30.635.2
Released2024-05-292025-03-17
WeightsProprietaryOpen
Context window256K262K
Max output8K16K
Input $ / M tokens$0.30$2
Output $ / M tokens$0.90$5
Results tracked722

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Medium leads

Codestral: 27.3 (#321), Magistral Medium: 39.1 (#161)

Coding benchmarks
BenchmarkCodestralMagistral Medium
Aider Polyglot11.1%—
SciCode—39.2%
BigCodeBench Instruct41.8%—
LMArena Coding—1319
BigCodeBench Complete52.5%—
ALE-Bench137.78—
HumanEval+73.8%—
MBPP+61.9%—

Reasoning Codestral leads

Codestral: 19.8 (#251), Magistral Medium: 8.6 (#348)

Reasoning benchmarks
BenchmarkCodestralMagistral Medium
Kagi LLM Benchmark32.5%16.2%
ARC-AGI-2—0%
ARC-AGI-1—6.1%
CritPt—0.3%
LMArena Hard Prompts—1267

Math Not comparable

Codestral: —, Magistral Medium: 35.1 (#189)

Math benchmarks
BenchmarkCodestralMagistral Medium
LMArena Math—1250

Knowledge Not comparable

Codestral: —, Magistral Medium: 33.5 (#202)

Knowledge benchmarks
BenchmarkCodestralMagistral Medium
LMArena Expert—1223

Multilingual Not comparable

Codestral: —, Magistral Medium: 39.6 (#224)

Multilingual benchmarks
BenchmarkCodestralMagistral Medium
LMArena Non-English—1232
LMArena Chinese—1227
LMArena French—1267
LMArena German—1248
LMArena Japanese—1175
LMArena Korean—1125
LMArena Russian—1224
LMArena Spanish—1271

Instruction Following Not comparable

Codestral: —, Magistral Medium: 66.0 (#211)

Instruction Following benchmarks
BenchmarkCodestralMagistral Medium
LMArena Instruction Following—1254

Long Context Not comparable

Codestral: —, Magistral Medium: 39.3 (#183)

Long Context benchmarks
BenchmarkCodestralMagistral Medium
LMArena Longer Query—1295

Writing & Preference Not comparable

Codestral: —, Magistral Medium: 46.3 (#219)

Writing & Preference benchmarks
BenchmarkCodestralMagistral Medium
LMArena Text—1255
LMArena Creative Writing—1245
LMArena Multi-Turn—1275

Frequently asked questions

Is Codestral better than Magistral Medium?

Magistral Medium is the stronger model overall, scoring 35.2 to 30.6 on the Noometry Index. Codestral costs 6.1× less per token, which makes it the better buy when Magistral Medium's lead doesn't matter for your workload.

Which is cheaper, Codestral or Magistral Medium?

Codestral is cheaper. It lists at $0.30 per million input tokens and $0.90 per million output tokens; Magistral Medium lists at $2 and $5.

Is Codestral or Magistral Medium better for coding?

Magistral Medium scores higher on coding benchmarks: 39.1 versus 27.3 in the Noometry coding category.

Which has the bigger context window?

Magistral Medium does, with 262K tokens against 256K.

How many benchmarks do Codestral and Magistral Medium share?

1 benchmark has published results for both models. Codestral has 7 scored results on Noometry and Magistral Medium has 22.

Related comparisons

Go deeper