Model comparison

Mistral Large vs Mixtral 8x22B

Mistral Large is the stronger model overall, scoring 31.9 to 27.1 on the Noometry Index.

Last verified . 32 shared benchmarks.

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Mixtral 8x22B Mistral AI

27.1

Rank #333 Confirmed

Summary

  • They share 32 benchmarks with published results for both. Mistral Large scores higher in 7 categories and Mixtral 8x22B in 2 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Mistral Large leads 30.1 to 15.1.
  • The biggest single-benchmark swing is MATH Level 5: 50.3% for Mistral Large and 24.2% for Mixtral 8x22B.
  • Both cost about the same: $2 input and $6 output per million tokens.
  • Mistral Large accepts more context: 131K tokens versus 64K.

Side by side

Mistral Large and Mixtral 8x22B specifications
Mistral LargeMixtral 8x22B
ProviderMistral AIMistral AI
Noometry Index31.927.1
Released2024-02-262024-04-17
WeightsOpenOpen
Context window131K64K
Max output16K64K
Input $ / M tokens$2$2
Output $ / M tokens$6$6
Results tracked5134

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Large leads

Mistral Large: 34.3 (#240), Mixtral 8x22B: 24.2 (#329)

Coding benchmarks
BenchmarkMistral LargeMixtral 8x22B
BigCodeBench Instruct30%40.6%
LMArena Coding12771166
BigCodeBench Complete38.3%50.2%
HumanEval+62.2%72%
MBPP+59.5%64.3%
SciCode36.2%—
WeirdML—3.2%
LiveBench Coding47.1%—
ALE-Bench264.7—

Agentic & Tool Use Mistral Large leads

Mistral Large: 28.6 (#89), Mixtral 8x22B: 23.1 (#127)

Agentic & Tool Use benchmarks
BenchmarkMistral LargeMixtral 8x22B
Berkeley Function Calling Leaderboard38.4%—
Cybench—7.5%

Reasoning Mixtral 8x22B leads

Mistral Large: 15.8 (#310), Mixtral 8x22B: 19.9 (#248)

Reasoning benchmarks
BenchmarkMistral LargeMixtral 8x22B
LMArena Hard Prompts12571150
DTBench65.1%55.1%
Epoch Capabilities Index128.52122.03
ForecastBench57.156.3
SimpleBench22.5%—
CritPt0%—
LiveBench Reasoning43.5%—
LiveBench Data Analysis50.1%—
LMCA16.7%—
LiveBench48.4%—

Math Mixtral 8x22B leads

Mistral Large: 18.2 (#291), Mixtral 8x22B: 22.9 (#275)

Math benchmarks
BenchmarkMistral LargeMixtral 8x22B
Omni-MATH28.1%16.3%
LMArena Math12621184
MATH Level 550.3%24.2%
OTIS Mock AIME 2024-20258.5%—
LiveBench Math42.5%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Mistral Large leads

Mistral Large: 30.1 (#230), Mixtral 8x22B: 15.1 (#293)

Knowledge benchmarks
BenchmarkMistral LargeMixtral 8x22B
GPQA Diamond51.3%34.1%
MMLU-Pro59.9%46%
GPQA (HELM)43.5%33.4%
LMArena Expert12321113
MMLU80%77.8%
Confabulations21.4%—
Vectara Hallucination Rate4.5%—

Multilingual Mistral Large leads

Mistral Large: 40.0 (#219), Mixtral 8x22B: 32.8 (#255)

Multilingual benchmarks
BenchmarkMistral LargeMixtral 8x22B
LMArena Non-English12371128
LMArena Chinese12401116
LMArena French13251166
LMArena German12541141
LMArena Japanese11881037
LMArena Korean12021057
LMArena Russian12571158
LMArena Spanish12681151

Instruction Following Mistral Large leads

Mistral Large: 67.9 (#191), Mixtral 8x22B: 57.7 (#266)

Instruction Following benchmarks
BenchmarkMistral LargeMixtral 8x22B
IFEval87.7%72.4%
LMArena Instruction Following12491147
LiveBench Instruction Following67.9%—

Long Context Mistral Large leads

Mistral Large: 38.3 (#199), Mixtral 8x22B: 34.7 (#247)

Long Context benchmarks
BenchmarkMistral LargeMixtral 8x22B
LMArena Longer Query12611144

Writing & Preference Mistral Large leads

Mistral Large: 40.7 (#242), Mixtral 8x22B: 36.9 (#262)

Writing & Preference benchmarks
BenchmarkMistral LargeMixtral 8x22B
LMArena Text12661162
LMArena Creative Writing12431141
WildBench80.1%71.1%
LMArena Multi-Turn12601130
Short-Story Creative Writing69%—
EQ-Bench Creative Writing985—
LiveBench Language39.4%—

Frequently asked questions

Is Mistral Large better than Mixtral 8x22B?

Mistral Large is the stronger model overall, scoring 31.9 to 27.1 on the Noometry Index.

Which is cheaper, Mistral Large or Mixtral 8x22B?

Mixtral 8x22B is cheaper. It lists at $2 per million input tokens and $6 per million output tokens; Mistral Large lists at $2 and $6.

Is Mistral Large or Mixtral 8x22B better for coding?

Mistral Large scores higher on coding benchmarks: 34.3 versus 24.2 in the Noometry coding category.

Which has the bigger context window?

Mistral Large does, with 131K tokens against 64K.

How many benchmarks do Mistral Large and Mixtral 8x22B share?

32 benchmarks have published results for both models. Mistral Large has 51 scored results on Noometry and Mixtral 8x22B has 34.

Related comparisons

Go deeper