Model comparison

Mistral Medium 3.1 vs Mixtral 8x7B

Mistral Medium 3.1 is the stronger model overall, scoring 31.9 to 27.1 on the Noometry Index.

Last verified . 0 shared benchmarks.

Mistral Medium 3.1 Mistral AI

31.9

Rank #266 Reported

Mixtral 8x7B Mistral AI

27.1

Rank #334 Confirmed

Summary

  • The widest gap is in writing & preference, where Mistral Medium 3.1 leads 55.5 to 34.2.
  • Mixtral 8x7B is cheaper at $0.70 / $0.70 per million input/output tokens, against $0.40 / $2 for Mistral Medium 3.1.
  • Mistral Medium 3.1 accepts more context: 131K tokens versus 32K.
  • Mixtral 8x7B has downloadable open weights; the other is API-only.

Side by side

Mistral Medium 3.1 and Mixtral 8x7B specifications
Mistral Medium 3.1Mixtral 8x7B
ProviderMistral AIMistral AI
Noometry Index31.927.1
Released—2023-12-11
WeightsProprietaryOpen
Context window131K32K
Max output105K32K
Input $ / M tokens$0.40$0.70
Output $ / M tokens$2$0.70
Results tracked338

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Mistral Medium 3.1: —, Mixtral 8x7B: 32.8 (#269)

Coding benchmarks
BenchmarkMistral Medium 3.1Mixtral 8x7B
LMArena Coding—1126
HumanEval+—39.6%
MBPP+—49.7%

Reasoning Mixtral 8x7B leads

Mistral Medium 3.1: 10.6 (#341), Mixtral 8x7B: 18.2 (#285)

Reasoning benchmarks
BenchmarkMistral Medium 3.1Mixtral 8x7B
NYT Connections (extended)6.5%—
Thematic Generalization20.3%—
LMArena Hard Prompts—1115
DTBench—49.6%
Adversarial NLI—55.2%
Epoch Capabilities Index—118.47
ForecastBench—56.3
HellaSwag—86.7%
PIQA—83.6%
WinoGrande—77.2%

Math Not comparable

Mistral Medium 3.1: —, Mixtral 8x7B: 18.8 (#289)

Math benchmarks
BenchmarkMistral Medium 3.1Mixtral 8x7B
Omni-MATH—10.5%
LMArena Math—1147
MATH Level 5—10%
GSM8K—74.4%

Knowledge Not comparable

Mistral Medium 3.1: —, Mixtral 8x7B: 11.0 (#301)

Knowledge benchmarks
BenchmarkMistral Medium 3.1Mixtral 8x7B
GPQA Diamond—30.6%
MMLU-Pro—33.5%
GPQA (HELM)—29.6%
LMArena Expert—1088
ARC (AI2) Challenge—87.3%
MMLU—70.6%
OpenBookQA—85.8%
TriviaQA—82.2%

Multilingual Not comparable

Mistral Medium 3.1: —, Mixtral 8x7B: 29.6 (#266)

Multilingual benchmarks
BenchmarkMistral Medium 3.1Mixtral 8x7B
LMArena Non-English—1077
LMArena Chinese—1055
LMArena French—1166
LMArena German—1114
LMArena Japanese—931
LMArena Korean—968
LMArena Russian—1090
LMArena Spanish—1111

Instruction Following Not comparable

Mistral Medium 3.1: —, Mixtral 8x7B: 51.0 (#297)

Instruction Following benchmarks
BenchmarkMistral Medium 3.1Mixtral 8x7B
IFEval—57.5%
LMArena Instruction Following—1109

Long Context Not comparable

Mistral Medium 3.1: —, Mixtral 8x7B: 33.4 (#260)

Long Context benchmarks
BenchmarkMistral Medium 3.1Mixtral 8x7B
LMArena Longer Query—1103

Writing & Preference Mistral Medium 3.1 leads

Mistral Medium 3.1: 55.5 (#145), Mixtral 8x7B: 34.2 (#270)

Writing & Preference benchmarks
BenchmarkMistral Medium 3.1Mixtral 8x7B
LMArena Text—1132
LMArena Creative Writing—1109
EQ-Bench Creative Writing1476—
WildBench—67.3%
LMArena Multi-Turn—1115

Frequently asked questions

Is Mistral Medium 3.1 better than Mixtral 8x7B?

Mistral Medium 3.1 is the stronger model overall, scoring 31.9 to 27.1 on the Noometry Index.

Which is cheaper, Mistral Medium 3.1 or Mixtral 8x7B?

Mixtral 8x7B is cheaper. It lists at $0.70 per million input tokens and $0.70 per million output tokens; Mistral Medium 3.1 lists at $0.40 and $2.

Which has the bigger context window?

Mistral Medium 3.1 does, with 131K tokens against 32K.

How many benchmarks do Mistral Medium 3.1 and Mixtral 8x7B share?

0 benchmarks have published results for both models. Mistral Medium 3.1 has 3 scored results on Noometry and Mixtral 8x7B has 38.

Related comparisons

Go deeper