Model comparison

Mistral vs Mixtral 8x22B

Mistral is the stronger model overall, scoring 29.9 to 27.1 on the Noometry Index.

Last verified . 22 shared benchmarks.

Mistral Mistral AI

29.9

Rank #303 Confirmed

Mixtral 8x22B Mistral AI

27.1

Rank #333 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Mistral scores higher in 6 categories and Mixtral 8x22B in 2 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Mistral leads 33.8 to 24.2.
  • The biggest single-benchmark swing is MMLU-Pro: 27.7% for Mistral and 46% for Mixtral 8x22B.
  • Mixtral 8x22B has downloadable open weights; the other is API-only.

Side by side

Mistral and Mixtral 8x22B specifications
MistralMixtral 8x22B
ProviderMistral AIMistral AI
Noometry Index29.927.1
Released—2024-04-17
WeightsProprietaryOpen
Context window—64K
Max output—64K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked2234

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral leads

Mistral: 33.8 (#250), Mixtral 8x22B: 24.2 (#329)

Coding benchmarks
BenchmarkMistralMixtral 8x22B
LMArena Coding11621166
WeirdML—3.2%
BigCodeBench Instruct—40.6%
BigCodeBench Complete—50.2%
HumanEval+—72%
MBPP+—64.3%

Agentic & Tool Use Not comparable

Mistral: —, Mixtral 8x22B: 23.1 (#127)

Agentic & Tool Use benchmarks
BenchmarkMistralMixtral 8x22B
Cybench—7.5%

Reasoning Mistral leads

Mistral: 22.2 (#200), Mixtral 8x22B: 19.9 (#248)

Reasoning benchmarks
BenchmarkMistralMixtral 8x22B
LMArena Hard Prompts11491150
DTBench—55.1%
Epoch Capabilities Index—122.03
ForecastBench—56.3

Math Too close to call

Mistral: 22.3 (#278), Mixtral 8x22B: 22.9 (#275)

Math benchmarks
BenchmarkMistralMixtral 8x22B
Omni-MATH7.2%16.3%
LMArena Math11801184
MATH Level 5—24.2%

Knowledge Mistral leads

Mistral: 16.6 (#288), Mixtral 8x22B: 15.1 (#293)

Knowledge benchmarks
BenchmarkMistralMixtral 8x22B
MMLU-Pro27.7%46%
GPQA (HELM)30.3%33.4%
LMArena Expert11251113
GPQA Diamond—34.1%
MMLU—77.8%

Multilingual Too close to call

Mistral: 32.8 (#254), Mixtral 8x22B: 32.8 (#255)

Multilingual benchmarks
BenchmarkMistralMixtral 8x22B
LMArena Non-English11291128
LMArena Chinese11091116
LMArena French11801166
LMArena German11551141
LMArena Japanese10131037
LMArena Korean10321057
LMArena Russian11681158
LMArena Spanish11431151

Instruction Following Mixtral 8x22B leads

Mistral: 52.6 (#288), Mixtral 8x22B: 57.7 (#266)

Instruction Following benchmarks
BenchmarkMistralMixtral 8x22B
IFEval56.8%72.4%
LMArena Instruction Following11521147

Long Context Too close to call

Mistral: 35.0 (#245), Mixtral 8x22B: 34.7 (#247)

Long Context benchmarks
BenchmarkMistralMixtral 8x22B
LMArena Longer Query11531144

Writing & Preference Too close to call

Mistral: 37.0 (#260), Mixtral 8x22B: 36.9 (#262)

Writing & Preference benchmarks
BenchmarkMistralMixtral 8x22B
LMArena Text11651162
LMArena Creative Writing11581141
WildBench66%71.1%
LMArena Multi-Turn11471130

Frequently asked questions

Is Mistral better than Mixtral 8x22B?

Mistral is the stronger model overall, scoring 29.9 to 27.1 on the Noometry Index.

Is Mistral or Mixtral 8x22B better for coding?

Mistral scores higher on coding benchmarks: 33.8 versus 24.2 in the Noometry coding category.

How many benchmarks do Mistral and Mixtral 8x22B share?

22 benchmarks have published results for both models. Mistral has 22 scored results on Noometry and Mixtral 8x22B has 34.

Related comparisons

Go deeper