Model comparison

Mistral Medium 3.5 vs Mixtral 8x22B

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 27.1 on the Noometry Index.

Last verified . 17 shared benchmarks.

Mistral Medium 3.5 Mistral AI

40.2

Rank #152 Confirmed

Mixtral 8x22B Mistral AI

27.1

Rank #333 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Mistral Medium 3.5 scores higher in 7 categories and Mixtral 8x22B in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Mistral Medium 3.5 leads 40.0 to 15.1.
  • Both cost about the same: $1.50 input and $7.50 output per million tokens.
  • Mistral Medium 3.5 accepts more context: 262K tokens versus 64K.

Side by side

Mistral Medium 3.5 and Mixtral 8x22B specifications
Mistral Medium 3.5Mixtral 8x22B
ProviderMistral AIMistral AI
Noometry Index40.227.1
Released—2024-04-17
WeightsOpenOpen
Context window262K64K
Max output210K64K
Input $ / M tokens$1.50$2
Output $ / M tokens$7.50$6
Results tracked2234

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Medium 3.5 leads

Mistral Medium 3.5: 36.0 (#213), Mixtral 8x22B: 24.2 (#329)

Coding benchmarks
BenchmarkMistral Medium 3.5Mixtral 8x22B
LMArena Coding14611166
LMArena WebDev1264—
WeirdML—3.2%
BigCodeBench Instruct—40.6%
BigCodeBench Complete—50.2%
HumanEval+—72%
MBPP+—64.3%

Agentic & Tool Use Not comparable

Mistral Medium 3.5: —, Mixtral 8x22B: 23.1 (#127)

Agentic & Tool Use benchmarks
BenchmarkMistral Medium 3.5Mixtral 8x22B
Cybench—7.5%

Reasoning Mixtral 8x22B leads

Mistral Medium 3.5: 17.3 (#295), Mixtral 8x22B: 19.9 (#248)

Reasoning benchmarks
BenchmarkMistral Medium 3.5Mixtral 8x22B
LMArena Hard Prompts14361150
Epoch Capabilities Index141.35122.03
Kagi LLM Benchmark41.4%—
NYT Connections (extended)12.9%—
DTBench—55.1%
ForecastBench—56.3

Math Mistral Medium 3.5 leads

Mistral Medium 3.5: 39.1 (#113), Mixtral 8x22B: 22.9 (#275)

Math benchmarks
BenchmarkMistral Medium 3.5Mixtral 8x22B
LMArena Math14311184
Omni-MATH—16.3%
MATH Level 5—24.2%

Knowledge Mistral Medium 3.5 leads

Mistral Medium 3.5: 40.0 (#126), Mixtral 8x22B: 15.1 (#293)

Knowledge benchmarks
BenchmarkMistral Medium 3.5Mixtral 8x22B
LMArena Expert14321113
GPQA Diamond—34.1%
MMLU-Pro—46%
GPQA (HELM)—33.4%
MMLU—77.8%

Multimodal Not comparable

Mistral Medium 3.5: 38.3 (#65), Mixtral 8x22B: —

Multimodal benchmarks
BenchmarkMistral Medium 3.5Mixtral 8x22B
LMArena Vision1223—

Multilingual Mistral Medium 3.5 leads

Mistral Medium 3.5: 51.9 (#100), Mixtral 8x22B: 32.8 (#255)

Multilingual benchmarks
BenchmarkMistral Medium 3.5Mixtral 8x22B
LMArena Non-English14041128
LMArena Chinese14421116
LMArena French14481166
LMArena German14511141
LMArena Korean13851057
LMArena Russian13951158
LMArena Spanish14091151
LMArena Japanese—1037

Instruction Following Mistral Medium 3.5 leads

Mistral Medium 3.5: 74.6 (#90), Mixtral 8x22B: 57.7 (#266)

Instruction Following benchmarks
BenchmarkMistral Medium 3.5Mixtral 8x22B
LMArena Instruction Following14151147
IFEval—72.4%

Long Context Mistral Medium 3.5 leads

Mistral Medium 3.5: 43.2 (#103), Mixtral 8x22B: 34.7 (#247)

Long Context benchmarks
BenchmarkMistral Medium 3.5Mixtral 8x22B
LMArena Longer Query14151144

Writing & Preference Mistral Medium 3.5 leads

Mistral Medium 3.5: 58.5 (#117), Mixtral 8x22B: 36.9 (#262)

Writing & Preference benchmarks
BenchmarkMistral Medium 3.5Mixtral 8x22B
LMArena Text14211162
LMArena Creative Writing13741141
LMArena Multi-Turn14231130
WildBench—71.1%
EQ-Bench 4993—

Frequently asked questions

Is Mistral Medium 3.5 better than Mixtral 8x22B?

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 27.1 on the Noometry Index.

Which is cheaper, Mistral Medium 3.5 or Mixtral 8x22B?

Mixtral 8x22B is cheaper. It lists at $2 per million input tokens and $6 per million output tokens; Mistral Medium 3.5 lists at $1.50 and $7.50.

Is Mistral Medium 3.5 or Mixtral 8x22B better for coding?

Mistral Medium 3.5 scores higher on coding benchmarks: 36.0 versus 24.2 in the Noometry coding category.

Which has the bigger context window?

Mistral Medium 3.5 does, with 262K tokens against 64K.

How many benchmarks do Mistral Medium 3.5 and Mixtral 8x22B share?

17 benchmarks have published results for both models. Mistral Medium 3.5 has 22 scored results on Noometry and Mixtral 8x22B has 34.

Related comparisons

Go deeper