Model comparison

Mercury 2.5 vs Mistral Medium

Mistral Medium is the stronger model overall, scoring 36.3 to 33.5 on the Noometry Index. Mercury 2.5 costs 44× less per token, which makes it the better buy when Mistral Medium's lead doesn't matter for your workload.

Last verified . 4 shared benchmarks.

Mercury 2.5 Inception

33.5

Rank #242 Reported

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Mercury 2.5 scores higher in 1 category and Mistral Medium in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Mercury 2.5 leads 39.5 to 34.2.
  • The biggest single-benchmark swing is ProofBench: 3% for Mercury 2.5 and 9% for Mistral Medium.
  • Mercury 2.5 is cheaper at $0.04 / $0.15 per million input/output tokens, against $1.50 / $7.50 for Mistral Medium.
  • Mistral Medium accepts more context: 262K tokens versus 260K.
  • Mistral Medium has downloadable open weights; the other is API-only.

Side by side

Mercury 2.5 and Mistral Medium specifications
Mercury 2.5Mistral Medium
ProviderInceptionMistral AI
Noometry Index33.536.3
Released2026-09-082023-12-11
WeightsProprietaryOpen
Context window260K262K
Max output66K262K
Input $ / M tokens$0.04$1.50
Output $ / M tokens$0.15$7.50
Results tracked436

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury 2.5 leads

Mercury 2.5: 39.5 (#156), Mistral Medium: 34.2 (#243)

Coding benchmarks
BenchmarkMercury 2.5Mistral Medium
SciCode38.5%40.2%
ALE-Bench301.65763.98
FrontierCode—8%
WeirdML—43.7%
LMArena Coding—1434

Agentic & Tool Use Not comparable

Mercury 2.5: —, Mistral Medium: 28.3 (#90)

Agentic & Tool Use benchmarks
BenchmarkMercury 2.5Mistral Medium
Berkeley Function Calling Leaderboard—37.7%

Reasoning Mistral Medium leads

Mercury 2.5: 22.4 (#193), Mistral Medium: 24.0 (#167)

Reasoning benchmarks
BenchmarkMercury 2.5Mistral Medium
CritPt0%0%
Kagi LLM Benchmark—50%
LMArena Hard Prompts—1426
DTBench—75.5%
LMCA—26.1%
Surface Evolver Bench—26.9%

Math Mistral Medium leads

Mercury 2.5: 23.3 (#272), Mistral Medium: 28.1 (#245)

Math benchmarks
BenchmarkMercury 2.5Mistral Medium
ProofBench3%9%
OTIS Mock AIME 2024-2025—32.2%
LMArena Math—1408
MATH Level 5—81.6%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Not comparable

Mercury 2.5: —, Mistral Medium: 25.0 (#265)

Knowledge benchmarks
BenchmarkMercury 2.5Mistral Medium
GPQA Diamond—59.5%
Humanity's Last Exam—4.5%
Vectara Hallucination Rate—22.7%
LMArena Expert—1408

Multimodal Not comparable

Mercury 2.5: —, Mistral Medium: 35.3 (#88)

Multimodal benchmarks
BenchmarkMercury 2.5Mistral Medium
LMArena Vision—1172

Multilingual Not comparable

Mercury 2.5: —, Mistral Medium: 52.1 (#91)

Multilingual benchmarks
BenchmarkMercury 2.5Mistral Medium
LMArena Non-English—1408
LMArena Chinese—1447
LMArena French—1459
LMArena German—1432
LMArena Japanese—1378
LMArena Korean—1380
LMArena Russian—1411
LMArena Spanish—1433

Instruction Following Not comparable

Mercury 2.5: —, Mistral Medium: 73.7 (#116)

Instruction Following benchmarks
BenchmarkMercury 2.5Mistral Medium
LMArena Instruction Following—1398

Long Context Not comparable

Mercury 2.5: —, Mistral Medium: 42.9 (#114)

Long Context benchmarks
BenchmarkMercury 2.5Mistral Medium
LMArena Longer Query—1406

Writing & Preference Not comparable

Mercury 2.5: —, Mistral Medium: 60.0 (#103)

Writing & Preference benchmarks
BenchmarkMercury 2.5Mistral Medium
LMArena Text—1424
LMArena Creative Writing—1391
Short-Story Creative Writing—77.3%
LMArena Multi-Turn—1418

Frequently asked questions

Is Mercury 2.5 better than Mistral Medium?

Mistral Medium is the stronger model overall, scoring 36.3 to 33.5 on the Noometry Index. Mercury 2.5 costs 44× less per token, which makes it the better buy when Mistral Medium's lead doesn't matter for your workload.

Which is cheaper, Mercury 2.5 or Mistral Medium?

Mercury 2.5 is cheaper. It lists at $0.04 per million input tokens and $0.15 per million output tokens; Mistral Medium lists at $1.50 and $7.50.

Is Mercury 2.5 or Mistral Medium better for coding?

Mercury 2.5 scores higher on coding benchmarks: 39.5 versus 34.2 in the Noometry coding category.

Which has the bigger context window?

Mistral Medium does, with 262K tokens against 260K.

How many benchmarks do Mercury 2.5 and Mistral Medium share?

4 benchmarks have published results for both models. Mercury 2.5 has 4 scored results on Noometry and Mistral Medium has 36.

Related comparisons

Go deeper