Model comparison

Mercury vs Mistral Large

Mercury is the stronger model overall, scoring 37.6 to 31.9 on the Noometry Index.

Last verified . 8 shared benchmarks.

Mercury Inception

37.6

Rank #199 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Mercury scores higher in 5 categories and Mistral Large in 1 category; 5 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mercury leads 46.2 to 40.7.
  • Mistral Large has downloadable open weights; the other is API-only.

Side by side

Mercury and Mistral Large specifications
MercuryMistral Large
ProviderInceptionMistral AI
Noometry Index37.631.9
Released—2024-02-26
WeightsProprietaryOpen
Context window—131K
Max output—16K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked951

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury leads

Mercury: 38.7 (#170), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkMercuryMistral Large
LMArena Coding13221277
SciCode—36.2%
BigCodeBench Instruct—30%
LiveBench Coding—47.1%
BigCodeBench Complete—38.3%
ALE-Bench—264.7
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Not comparable

Mercury: —, Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkMercuryMistral Large
Berkeley Function Calling Leaderboard—38.4%

Reasoning Mercury leads

Mercury: 17.5 (#293), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkMercuryMistral Large
LMArena Hard Prompts12851257
SimpleBench—22.5%
Kagi LLM Benchmark21.6%—
CritPt—0%
LiveBench Reasoning—43.5%
DTBench—65.1%
LiveBench Data Analysis—50.1%
LMCA—16.7%
Epoch Capabilities Index—128.52
ForecastBench—57.1
LiveBench—48.4%

Math Not comparable

Mercury: —, Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkMercuryMistral Large
OTIS Mock AIME 2024-2025—8.5%
Omni-MATH—28.1%
LiveBench Math—42.5%
LMArena Math—1262
MATH Level 5—50.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Not comparable

Mercury: —, Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkMercuryMistral Large
GPQA Diamond—51.3%
MMLU-Pro—59.9%
Confabulations—21.4%
Vectara Hallucination Rate—4.5%
GPQA (HELM)—43.5%
LMArena Expert—1232
MMLU—80%

Multilingual Mercury leads

Mercury: 41.6 (#206), Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkMercuryMistral Large
LMArena Non-English12601237
LMArena Chinese—1240
LMArena French—1325
LMArena German—1254
LMArena Japanese—1188
LMArena Korean—1202
LMArena Russian—1257
LMArena Spanish—1268

Instruction Following Mistral Large leads

Mercury: 65.2 (#224), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkMercuryMistral Large
LMArena Instruction Following12391249
LiveBench Instruction Following—67.9%
IFEval—87.7%

Long Context Too close to call

Mercury: 38.4 (#198), Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkMercuryMistral Large
LMArena Longer Query12661261

Writing & Preference Mercury leads

Mercury: 46.2 (#221), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkMercuryMistral Large
LMArena Text12821266
LMArena Creative Writing11911243
LMArena Multi-Turn12821260
Short-Story Creative Writing—69%
EQ-Bench Creative Writing—985
WildBench—80.1%
LiveBench Language—39.4%

Frequently asked questions

Is Mercury better than Mistral Large?

Mercury is the stronger model overall, scoring 37.6 to 31.9 on the Noometry Index.

Is Mercury or Mistral Large better for coding?

Mercury scores higher on coding benchmarks: 38.7 versus 34.3 in the Noometry coding category.

How many benchmarks do Mercury and Mistral Large share?

8 benchmarks have published results for both models. Mercury has 9 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper