Model comparison

Mistral Medium 3.1 vs Mixtral 8x22B

Mistral Medium 3.1 is the stronger model overall, scoring 31.9 to 27.1 on the Noometry Index.

Last verified . 0 shared benchmarks.

Mistral Medium 3.1 Mistral AI

31.9

Rank #266 Reported

Mixtral 8x22B Mistral AI

27.1

Rank #333 Confirmed

Summary

  • The widest gap is in writing & preference, where Mistral Medium 3.1 leads 55.5 to 36.9.
  • Mistral Medium 3.1 is cheaper at $0.40 / $2 per million input/output tokens, against $2 / $6 for Mixtral 8x22B.
  • Mistral Medium 3.1 accepts more context: 131K tokens versus 64K.
  • Mixtral 8x22B has downloadable open weights; the other is API-only.

Side by side

Mistral Medium 3.1 and Mixtral 8x22B specifications
Mistral Medium 3.1Mixtral 8x22B
ProviderMistral AIMistral AI
Noometry Index31.927.1
Released—2024-04-17
WeightsProprietaryOpen
Context window131K64K
Max output105K64K
Input $ / M tokens$0.40$2
Output $ / M tokens$2$6
Results tracked334

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Mistral Medium 3.1: —, Mixtral 8x22B: 24.2 (#329)

Coding benchmarks
BenchmarkMistral Medium 3.1Mixtral 8x22B
WeirdML—3.2%
BigCodeBench Instruct—40.6%
LMArena Coding—1166
BigCodeBench Complete—50.2%
HumanEval+—72%
MBPP+—64.3%

Agentic & Tool Use Not comparable

Mistral Medium 3.1: —, Mixtral 8x22B: 23.1 (#127)

Agentic & Tool Use benchmarks
BenchmarkMistral Medium 3.1Mixtral 8x22B
Cybench—7.5%

Reasoning Mixtral 8x22B leads

Mistral Medium 3.1: 10.6 (#341), Mixtral 8x22B: 19.9 (#248)

Reasoning benchmarks
BenchmarkMistral Medium 3.1Mixtral 8x22B
NYT Connections (extended)6.5%—
Thematic Generalization20.3%—
LMArena Hard Prompts—1150
DTBench—55.1%
Epoch Capabilities Index—122.03
ForecastBench—56.3

Math Not comparable

Mistral Medium 3.1: —, Mixtral 8x22B: 22.9 (#275)

Math benchmarks
BenchmarkMistral Medium 3.1Mixtral 8x22B
Omni-MATH—16.3%
LMArena Math—1184
MATH Level 5—24.2%

Knowledge Not comparable

Mistral Medium 3.1: —, Mixtral 8x22B: 15.1 (#293)

Knowledge benchmarks
BenchmarkMistral Medium 3.1Mixtral 8x22B
GPQA Diamond—34.1%
MMLU-Pro—46%
GPQA (HELM)—33.4%
LMArena Expert—1113
MMLU—77.8%

Multilingual Not comparable

Mistral Medium 3.1: —, Mixtral 8x22B: 32.8 (#255)

Multilingual benchmarks
BenchmarkMistral Medium 3.1Mixtral 8x22B
LMArena Non-English—1128
LMArena Chinese—1116
LMArena French—1166
LMArena German—1141
LMArena Japanese—1037
LMArena Korean—1057
LMArena Russian—1158
LMArena Spanish—1151

Instruction Following Not comparable

Mistral Medium 3.1: —, Mixtral 8x22B: 57.7 (#266)

Instruction Following benchmarks
BenchmarkMistral Medium 3.1Mixtral 8x22B
IFEval—72.4%
LMArena Instruction Following—1147

Long Context Not comparable

Mistral Medium 3.1: —, Mixtral 8x22B: 34.7 (#247)

Long Context benchmarks
BenchmarkMistral Medium 3.1Mixtral 8x22B
LMArena Longer Query—1144

Writing & Preference Mistral Medium 3.1 leads

Mistral Medium 3.1: 55.5 (#145), Mixtral 8x22B: 36.9 (#262)

Writing & Preference benchmarks
BenchmarkMistral Medium 3.1Mixtral 8x22B
LMArena Text—1162
LMArena Creative Writing—1141
EQ-Bench Creative Writing1476—
WildBench—71.1%
LMArena Multi-Turn—1130

Frequently asked questions

Is Mistral Medium 3.1 better than Mixtral 8x22B?

Mistral Medium 3.1 is the stronger model overall, scoring 31.9 to 27.1 on the Noometry Index.

Which is cheaper, Mistral Medium 3.1 or Mixtral 8x22B?

Mistral Medium 3.1 is cheaper. It lists at $0.40 per million input tokens and $2 per million output tokens; Mixtral 8x22B lists at $2 and $6.

Which has the bigger context window?

Mistral Medium 3.1 does, with 131K tokens against 64K.

How many benchmarks do Mistral Medium 3.1 and Mixtral 8x22B share?

0 benchmarks have published results for both models. Mistral Medium 3.1 has 3 scored results on Noometry and Mixtral 8x22B has 34.

Related comparisons

Go deeper