Model comparison

Mercury vs Mistral 7B

Mercury is the stronger model overall, scoring 37.6 to 23.0 on the Noometry Index.

Last verified . 8 shared benchmarks.

Mercury Inception

37.6

Rank #199 Confirmed

Mistral 7B Mistral AI

23.0

Rank #351 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Mercury scores higher in 6 categories and Mistral 7B in 0 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Mercury leads 41.6 to 25.8.
  • Mistral 7B has downloadable open weights; the other is API-only.

Side by side

Mercury and Mistral 7B specifications
MercuryMistral 7B
ProviderInceptionMistral AI
Noometry Index37.623.0
Released—2023-09-27
WeightsProprietaryOpen
Context window—8K
Max output—8K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.25
Results tracked937

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury leads

Mercury: 38.7 (#170), Mistral 7B: 26.4 (#326)

Coding benchmarks
BenchmarkMercuryMistral 7B
LMArena Coding13221082
BigCodeBench Instruct—19.5%
BigCodeBench Complete—27.3%
HumanEval+—36%
MBPP+—42.1%

Reasoning Mercury leads

Mercury: 17.5 (#293), Mistral 7B: 13.1 (#336)

Reasoning benchmarks
BenchmarkMercuryMistral 7B
LMArena Hard Prompts12851067
Kagi LLM Benchmark21.6%—
Chess Puzzles—0%
DTBench—42.5%
Adversarial NLI—47.1%
BIG-Bench Hard—56.1%
Epoch Capabilities Index—112.21
HellaSwag—81%
PIQA—83%
WinoGrande—75.3%

Math Not comparable

Mercury: —, Mistral 7B: 8.1 (#325)

Math benchmarks
BenchmarkMercuryMistral 7B
OTIS Mock AIME 2024-2025—0.3%
LMArena Math—1085
MATH Level 5—3.7%
GSM8K—54.4%

Knowledge Not comparable

Mercury: —, Mistral 7B: 7.4 (#311)

Knowledge benchmarks
BenchmarkMercuryMistral 7B
GPQA Diamond—15.2%
LMArena Expert—1036
ARC (AI2) Challenge—78.6%
BoolQ—87.4%
MMLU—62.5%
OpenBookQA—79.8%
TriviaQA—75.2%

Multilingual Mercury leads

Mercury: 41.6 (#206), Mistral 7B: 25.8 (#283)

Multilingual benchmarks
BenchmarkMercuryMistral 7B
LMArena Non-English12601012
LMArena Chinese—1009
LMArena French—1037
LMArena German—987
LMArena Japanese—878
LMArena Russian—1018
LMArena Spanish—1026

Instruction Following Mercury leads

Mercury: 65.2 (#224), Mistral 7B: 54.2 (#280)

Instruction Following benchmarks
BenchmarkMercuryMistral 7B
LMArena Instruction Following12391060

Long Context Mercury leads

Mercury: 38.4 (#198), Mistral 7B: 32.2 (#271)

Long Context benchmarks
BenchmarkMercuryMistral 7B
LMArena Longer Query12661060

Writing & Preference Mercury leads

Mercury: 46.2 (#221), Mistral 7B: 30.7 (#286)

Writing & Preference benchmarks
BenchmarkMercuryMistral 7B
LMArena Text12821090
LMArena Creative Writing11911068
LMArena Multi-Turn12821062

Frequently asked questions

Is Mercury better than Mistral 7B?

Mercury is the stronger model overall, scoring 37.6 to 23.0 on the Noometry Index.

Is Mercury or Mistral 7B better for coding?

Mercury scores higher on coding benchmarks: 38.7 versus 26.4 in the Noometry coding category.

How many benchmarks do Mercury and Mistral 7B share?

8 benchmarks have published results for both models. Mercury has 9 scored results on Noometry and Mistral 7B has 37.

Related comparisons

Go deeper