Model comparison

Magistral Small vs Mercury

Mercury is the stronger model overall, scoring 37.6 to 30.2 on the Noometry Index.

Last verified . 1 shared benchmarks.

Magistral Small Mistral AI

30.2

Rank #296 Confirmed

Mercury Inception

37.6

Rank #199 Confirmed

Summary

  • They share 1 benchmark with published results for both. Magistral Small scores higher in 0 categories and Mercury in 2 categories; one gap is clear of the uncertainty.
  • The widest gap is in reasoning, where Mercury leads 17.5 to 6.8.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 6.3% for Magistral Small and 21.6% for Mercury.
  • Magistral Small has downloadable open weights; the other is API-only.

Side by side

Magistral Small and Mercury specifications
Magistral SmallMercury
ProviderMistral AIInception
Noometry Index30.237.6
Released2025-06-10—
WeightsOpenProprietary
Context window128K—
Max output40K—
Input $ / M tokens$0.50—
Output $ / M tokens$1.50—
Results tracked109

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Magistral Small: 38.4 (#176), Mercury: 38.7 (#170)

Coding benchmarks
BenchmarkMagistral SmallMercury
SciCode35.2%—
LMArena Coding—1322

Reasoning Mercury leads

Magistral Small: 6.8 (#350), Mercury: 17.5 (#293)

Reasoning benchmarks
BenchmarkMagistral SmallMercury
Kagi LLM Benchmark6.3%21.6%
ARC-AGI-20%—
ARC-AGI-15%—
CritPt0.3%—
Chess Puzzles3%—
LMArena Hard Prompts—1285
DTBench61.3%—
Epoch Capabilities Index133.19—

Math Not comparable

Magistral Small: 26.2 (#261), Mercury: —

Math benchmarks
BenchmarkMagistral SmallMercury
OTIS Mock AIME 2024-202530%—

Knowledge Not comparable

Magistral Small: 30.9 (#223), Mercury: —

Knowledge benchmarks
BenchmarkMagistral SmallMercury
GPQA Diamond56.1%—

Multilingual Not comparable

Magistral Small: —, Mercury: 41.6 (#206)

Multilingual benchmarks
BenchmarkMagistral SmallMercury
LMArena Non-English—1260

Instruction Following Not comparable

Magistral Small: —, Mercury: 65.2 (#224)

Instruction Following benchmarks
BenchmarkMagistral SmallMercury
LMArena Instruction Following—1239

Long Context Not comparable

Magistral Small: —, Mercury: 38.4 (#198)

Long Context benchmarks
BenchmarkMagistral SmallMercury
LMArena Longer Query—1266

Writing & Preference Not comparable

Magistral Small: —, Mercury: 46.2 (#221)

Writing & Preference benchmarks
BenchmarkMagistral SmallMercury
LMArena Text—1282
LMArena Creative Writing—1191
LMArena Multi-Turn—1282

Frequently asked questions

Is Magistral Small better than Mercury?

Mercury is the stronger model overall, scoring 37.6 to 30.2 on the Noometry Index.

Is Magistral Small or Mercury better for coding?

They score almost the same on coding (38.4 vs 38.7); test both on your own repository before choosing.

How many benchmarks do Magistral Small and Mercury share?

1 benchmark has published results for both models. Magistral Small has 10 scored results on Noometry and Mercury has 9.

Related comparisons

Go deeper