Model comparison

Magistral Small vs Mercury 2

Mercury 2 is the stronger model overall, scoring 39.1 to 30.2 on the Noometry Index.

Last verified . 2 shared benchmarks.

Magistral Small Mistral AI

30.2

Rank #296 Confirmed

Mercury 2 Inception

39.1

Rank #175 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Magistral Small scores higher in 1 category and Mercury 2 in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Mercury 2 leads 23.8 to 6.8.
  • Mercury 2 is cheaper at $0.25 / $0.75 per million input/output tokens, against $0.50 / $1.50 for Magistral Small.
  • Magistral Small has downloadable open weights; the other is API-only.

Side by side

Magistral Small and Mercury 2 specifications
Magistral SmallMercury 2
ProviderMistral AIInception
Noometry Index30.239.1
Released2025-06-102026-02-20
WeightsOpenProprietary
Context window128K128K
Max output40K50K
Input $ / M tokens$0.50$0.25
Output $ / M tokens$1.50$0.75
Results tracked1017

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Small leads

Magistral Small: 38.4 (#176), Mercury 2: 33.5 (#255)

Coding benchmarks
BenchmarkMagistral SmallMercury 2
SciCode35.2%38.7%
LMArena WebDev—1171
WeirdML—43.2%
LMArena Coding—1391
ALE-Bench—785.58

Reasoning Mercury 2 leads

Magistral Small: 6.8 (#350), Mercury 2: 23.8 (#170)

Reasoning benchmarks
BenchmarkMagistral SmallMercury 2
CritPt0.3%0.8%
ARC-AGI-20%—
Kagi LLM Benchmark6.3%—
ARC-AGI-15%—
Chess Puzzles3%—
LMArena Hard Prompts—1362
DTBench61.3%—
Epoch Capabilities Index133.19—

Math Not comparable

Magistral Small: 26.2 (#261), Mercury 2: —

Math benchmarks
BenchmarkMagistral SmallMercury 2
OTIS Mock AIME 2024-202530%—

Knowledge Mercury 2 leads

Magistral Small: 30.9 (#223), Mercury 2: 36.2 (#172)

Knowledge benchmarks
BenchmarkMagistral SmallMercury 2
GPQA Diamond56.1%—
Vectara Hallucination Rate—12.3%
LMArena Expert—1358

Multilingual Not comparable

Magistral Small: —, Mercury 2: 46.6 (#157)

Multilingual benchmarks
BenchmarkMagistral SmallMercury 2
LMArena Non-English—1331
LMArena Chinese—1417
LMArena Russian—1304

Instruction Following Not comparable

Magistral Small: —, Mercury 2: 70.2 (#165)

Instruction Following benchmarks
BenchmarkMagistral SmallMercury 2
LMArena Instruction Following—1329

Long Context Not comparable

Magistral Small: —, Mercury 2: 40.5 (#154)

Long Context benchmarks
BenchmarkMagistral SmallMercury 2
LMArena Longer Query—1330

Writing & Preference Not comparable

Magistral Small: —, Mercury 2: 53.8 (#155)

Writing & Preference benchmarks
BenchmarkMagistral SmallMercury 2
LMArena Text—1355
LMArena Creative Writing—1289
LMArena Multi-Turn—1358

Frequently asked questions

Is Magistral Small better than Mercury 2?

Mercury 2 is the stronger model overall, scoring 39.1 to 30.2 on the Noometry Index.

Which is cheaper, Magistral Small or Mercury 2?

Mercury 2 is cheaper. It lists at $0.25 per million input tokens and $0.75 per million output tokens; Magistral Small lists at $0.50 and $1.50.

Is Magistral Small or Mercury 2 better for coding?

Magistral Small scores higher on coding benchmarks: 38.4 versus 33.5 in the Noometry coding category.

Which has the bigger context window?

Both accept 128K tokens.

How many benchmarks do Magistral Small and Mercury 2 share?

2 benchmarks have published results for both models. Magistral Small has 10 scored results on Noometry and Mercury 2 has 17.

Related comparisons

Go deeper