Model comparison

Mercury 2.5 vs Mistral Large

Mercury 2.5 is the stronger model overall, scoring 33.5 to 31.9 on the Noometry Index.

Last verified . 3 shared benchmarks.

Mercury 2.5 Inception

33.5

Rank #242 Reported

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 3 benchmarks with published results for both. Mercury 2.5 scores higher in 3 categories and Mistral Large in 0 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Mercury 2.5 leads 22.4 to 15.8.
  • Mercury 2.5 is cheaper at $0.04 / $0.15 per million input/output tokens, against $2 / $6 for Mistral Large.
  • Mercury 2.5 accepts more context: 260K tokens versus 131K.
  • Mistral Large has downloadable open weights; the other is API-only.

Side by side

Mercury 2.5 and Mistral Large specifications
Mercury 2.5Mistral Large
ProviderInceptionMistral AI
Noometry Index33.531.9
Released2026-09-082024-02-26
WeightsProprietaryOpen
Context window260K131K
Max output66K16K
Input $ / M tokens$0.04$2
Output $ / M tokens$0.15$6
Results tracked451

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury 2.5 leads

Mercury 2.5: 39.5 (#156), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkMercury 2.5Mistral Large
SciCode38.5%36.2%
ALE-Bench301.65264.7
BigCodeBench Instruct—30%
LiveBench Coding—47.1%
LMArena Coding—1277
BigCodeBench Complete—38.3%
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Not comparable

Mercury 2.5: —, Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkMercury 2.5Mistral Large
Berkeley Function Calling Leaderboard—38.4%

Reasoning Mercury 2.5 leads

Mercury 2.5: 22.4 (#193), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkMercury 2.5Mistral Large
CritPt0%0%
SimpleBench—22.5%
LiveBench Reasoning—43.5%
LMArena Hard Prompts—1257
DTBench—65.1%
LiveBench Data Analysis—50.1%
LMCA—16.7%
Epoch Capabilities Index—128.52
ForecastBench—57.1
LiveBench—48.4%

Math Mercury 2.5 leads

Mercury 2.5: 23.3 (#272), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkMercury 2.5Mistral Large
OTIS Mock AIME 2024-2025—8.5%
ProofBench3%—
Omni-MATH—28.1%
LiveBench Math—42.5%
LMArena Math—1262
MATH Level 5—50.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Not comparable

Mercury 2.5: —, Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkMercury 2.5Mistral Large
GPQA Diamond—51.3%
MMLU-Pro—59.9%
Confabulations—21.4%
Vectara Hallucination Rate—4.5%
GPQA (HELM)—43.5%
LMArena Expert—1232
MMLU—80%

Multilingual Not comparable

Mercury 2.5: —, Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkMercury 2.5Mistral Large
LMArena Non-English—1237
LMArena Chinese—1240
LMArena French—1325
LMArena German—1254
LMArena Japanese—1188
LMArena Korean—1202
LMArena Russian—1257
LMArena Spanish—1268

Instruction Following Not comparable

Mercury 2.5: —, Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkMercury 2.5Mistral Large
LiveBench Instruction Following—67.9%
IFEval—87.7%
LMArena Instruction Following—1249

Long Context Not comparable

Mercury 2.5: —, Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkMercury 2.5Mistral Large
LMArena Longer Query—1261

Writing & Preference Not comparable

Mercury 2.5: —, Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkMercury 2.5Mistral Large
LMArena Text—1266
LMArena Creative Writing—1243
Short-Story Creative Writing—69%
EQ-Bench Creative Writing—985
WildBench—80.1%
LMArena Multi-Turn—1260
LiveBench Language—39.4%

Frequently asked questions

Is Mercury 2.5 better than Mistral Large?

Mercury 2.5 is the stronger model overall, scoring 33.5 to 31.9 on the Noometry Index.

Which is cheaper, Mercury 2.5 or Mistral Large?

Mercury 2.5 is cheaper. It lists at $0.04 per million input tokens and $0.15 per million output tokens; Mistral Large lists at $2 and $6.

Is Mercury 2.5 or Mistral Large better for coding?

Mercury 2.5 scores higher on coding benchmarks: 39.5 versus 34.3 in the Noometry coding category.

Which has the bigger context window?

Mercury 2.5 does, with 260K tokens against 131K.

How many benchmarks do Mercury 2.5 and Mistral Large share?

3 benchmarks have published results for both models. Mercury 2.5 has 4 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper