Model comparison

Mercury vs Mixtral 8x22B

Mercury is the stronger model overall, scoring 37.6 to 27.1 on the Noometry Index.

Last verified . 8 shared benchmarks.

Mercury Inception

37.6

Rank #199 Confirmed

Mixtral 8x22B Mistral AI

27.1

Rank #333 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Mercury scores higher in 5 categories and Mixtral 8x22B in 1 category; 6 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Mercury leads 38.7 to 24.2.
  • Mixtral 8x22B has downloadable open weights; the other is API-only.

Side by side

Mercury and Mixtral 8x22B specifications
MercuryMixtral 8x22B
ProviderInceptionMistral AI
Noometry Index37.627.1
Released—2024-04-17
WeightsProprietaryOpen
Context window—64K
Max output—64K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked934

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury leads

Mercury: 38.7 (#170), Mixtral 8x22B: 24.2 (#329)

Coding benchmarks
BenchmarkMercuryMixtral 8x22B
LMArena Coding13221166
WeirdML—3.2%
BigCodeBench Instruct—40.6%
BigCodeBench Complete—50.2%
HumanEval+—72%
MBPP+—64.3%

Agentic & Tool Use Not comparable

Mercury: —, Mixtral 8x22B: 23.1 (#127)

Agentic & Tool Use benchmarks
BenchmarkMercuryMixtral 8x22B
Cybench—7.5%

Reasoning Mixtral 8x22B leads

Mercury: 17.5 (#293), Mixtral 8x22B: 19.9 (#248)

Reasoning benchmarks
BenchmarkMercuryMixtral 8x22B
LMArena Hard Prompts12851150
Kagi LLM Benchmark21.6%—
DTBench—55.1%
Epoch Capabilities Index—122.03
ForecastBench—56.3

Math Not comparable

Mercury: —, Mixtral 8x22B: 22.9 (#275)

Math benchmarks
BenchmarkMercuryMixtral 8x22B
Omni-MATH—16.3%
LMArena Math—1184
MATH Level 5—24.2%

Knowledge Not comparable

Mercury: —, Mixtral 8x22B: 15.1 (#293)

Knowledge benchmarks
BenchmarkMercuryMixtral 8x22B
GPQA Diamond—34.1%
MMLU-Pro—46%
GPQA (HELM)—33.4%
LMArena Expert—1113
MMLU—77.8%

Multilingual Mercury leads

Mercury: 41.6 (#206), Mixtral 8x22B: 32.8 (#255)

Multilingual benchmarks
BenchmarkMercuryMixtral 8x22B
LMArena Non-English12601128
LMArena Chinese—1116
LMArena French—1166
LMArena German—1141
LMArena Japanese—1037
LMArena Korean—1057
LMArena Russian—1158
LMArena Spanish—1151

Instruction Following Mercury leads

Mercury: 65.2 (#224), Mixtral 8x22B: 57.7 (#266)

Instruction Following benchmarks
BenchmarkMercuryMixtral 8x22B
LMArena Instruction Following12391147
IFEval—72.4%

Long Context Mercury leads

Mercury: 38.4 (#198), Mixtral 8x22B: 34.7 (#247)

Long Context benchmarks
BenchmarkMercuryMixtral 8x22B
LMArena Longer Query12661144

Writing & Preference Mercury leads

Mercury: 46.2 (#221), Mixtral 8x22B: 36.9 (#262)

Writing & Preference benchmarks
BenchmarkMercuryMixtral 8x22B
LMArena Text12821162
LMArena Creative Writing11911141
LMArena Multi-Turn12821130
WildBench—71.1%

Frequently asked questions

Is Mercury better than Mixtral 8x22B?

Mercury is the stronger model overall, scoring 37.6 to 27.1 on the Noometry Index.

Is Mercury or Mixtral 8x22B better for coding?

Mercury scores higher on coding benchmarks: 38.7 versus 24.2 in the Noometry coding category.

How many benchmarks do Mercury and Mixtral 8x22B share?

8 benchmarks have published results for both models. Mercury has 9 scored results on Noometry and Mixtral 8x22B has 34.

Related comparisons

Go deeper