Model comparison

Magistral Small vs Mixtral 8x7B

Magistral Small is the stronger model overall, scoring 30.2 to 27.1 on the Noometry Index.

Last verified . 3 shared benchmarks.

Magistral Small Mistral AI

30.2

Rank #296 Confirmed

Mixtral 8x7B Mistral AI

27.1

Rank #334 Confirmed

Summary

  • They share 3 benchmarks with published results for both. Magistral Small scores higher in 3 categories and Mixtral 8x7B in 1 category; 4 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Magistral Small leads 30.9 to 11.0.
  • The biggest single-benchmark swing is GPQA Diamond: 56.1% for Magistral Small and 30.6% for Mixtral 8x7B.
  • Mixtral 8x7B is cheaper at $0.70 / $0.70 per million input/output tokens, against $0.50 / $1.50 for Magistral Small.
  • Magistral Small accepts more context: 128K tokens versus 32K.

Side by side

Magistral Small and Mixtral 8x7B specifications
Magistral SmallMixtral 8x7B
ProviderMistral AIMistral AI
Noometry Index30.227.1
Released2025-06-102023-12-11
WeightsOpenOpen
Context window128K32K
Max output40K32K
Input $ / M tokens$0.50$0.70
Output $ / M tokens$1.50$0.70
Results tracked1038

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Small leads

Magistral Small: 38.4 (#176), Mixtral 8x7B: 32.8 (#269)

Coding benchmarks
BenchmarkMagistral SmallMixtral 8x7B
SciCode35.2%—
LMArena Coding—1126
HumanEval+—39.6%
MBPP+—49.7%

Reasoning Mixtral 8x7B leads

Magistral Small: 6.8 (#350), Mixtral 8x7B: 18.2 (#285)

Reasoning benchmarks
BenchmarkMagistral SmallMixtral 8x7B
DTBench61.3%49.6%
Epoch Capabilities Index133.19118.47
ARC-AGI-20%—
Kagi LLM Benchmark6.3%—
ARC-AGI-15%—
CritPt0.3%—
Chess Puzzles3%—
LMArena Hard Prompts—1115
Adversarial NLI—55.2%
ForecastBench—56.3
HellaSwag—86.7%
PIQA—83.6%
WinoGrande—77.2%

Math Magistral Small leads

Magistral Small: 26.2 (#261), Mixtral 8x7B: 18.8 (#289)

Math benchmarks
BenchmarkMagistral SmallMixtral 8x7B
OTIS Mock AIME 2024-202530%—
Omni-MATH—10.5%
LMArena Math—1147
MATH Level 5—10%
GSM8K—74.4%

Knowledge Magistral Small leads

Magistral Small: 30.9 (#223), Mixtral 8x7B: 11.0 (#301)

Knowledge benchmarks
BenchmarkMagistral SmallMixtral 8x7B
GPQA Diamond56.1%30.6%
MMLU-Pro—33.5%
GPQA (HELM)—29.6%
LMArena Expert—1088
ARC (AI2) Challenge—87.3%
MMLU—70.6%
OpenBookQA—85.8%
TriviaQA—82.2%

Multilingual Not comparable

Magistral Small: —, Mixtral 8x7B: 29.6 (#266)

Multilingual benchmarks
BenchmarkMagistral SmallMixtral 8x7B
LMArena Non-English—1077
LMArena Chinese—1055
LMArena French—1166
LMArena German—1114
LMArena Japanese—931
LMArena Korean—968
LMArena Russian—1090
LMArena Spanish—1111

Instruction Following Not comparable

Magistral Small: —, Mixtral 8x7B: 51.0 (#297)

Instruction Following benchmarks
BenchmarkMagistral SmallMixtral 8x7B
IFEval—57.5%
LMArena Instruction Following—1109

Long Context Not comparable

Magistral Small: —, Mixtral 8x7B: 33.4 (#260)

Long Context benchmarks
BenchmarkMagistral SmallMixtral 8x7B
LMArena Longer Query—1103

Writing & Preference Not comparable

Magistral Small: —, Mixtral 8x7B: 34.2 (#270)

Writing & Preference benchmarks
BenchmarkMagistral SmallMixtral 8x7B
LMArena Text—1132
LMArena Creative Writing—1109
WildBench—67.3%
LMArena Multi-Turn—1115

Frequently asked questions

Is Magistral Small better than Mixtral 8x7B?

Magistral Small is the stronger model overall, scoring 30.2 to 27.1 on the Noometry Index.

Which is cheaper, Magistral Small or Mixtral 8x7B?

Mixtral 8x7B is cheaper. It lists at $0.70 per million input tokens and $0.70 per million output tokens; Magistral Small lists at $0.50 and $1.50.

Is Magistral Small or Mixtral 8x7B better for coding?

Magistral Small scores higher on coding benchmarks: 38.4 versus 32.8 in the Noometry coding category.

Which has the bigger context window?

Magistral Small does, with 128K tokens against 32K.

How many benchmarks do Magistral Small and Mixtral 8x7B share?

3 benchmarks have published results for both models. Magistral Small has 10 scored results on Noometry and Mixtral 8x7B has 38.

Related comparisons

Go deeper