Model comparison

Mixtral 8x7B vs Phi 3 Mini 128k Instruct

Phi 3 Mini 128k Instruct is the stronger model overall, scoring 29.7 to 27.1 on the Noometry Index.

Last verified . 17 shared benchmarks.

Mixtral 8x7B Mistral AI

27.1

Rank #334 Confirmed

Phi 3 Mini 128k Instruct Microsoft

29.7

Rank #305 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Mixtral 8x7B scores higher in 4 categories and Phi 3 Mini 128k Instruct in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Phi 3 Mini 128k Instruct leads 26.8 to 11.0.

Side by side

Mixtral 8x7B and Phi 3 Mini 128k Instruct specifications
Mixtral 8x7BPhi 3 Mini 128k Instruct
ProviderMistral AIMicrosoft
Noometry Index27.129.7
Released2023-12-112024-04-23
WeightsOpenOpen
Context window32K—
Max output32K—
Input $ / M tokens$0.70—
Output $ / M tokens$0.70—
Results tracked3819

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mixtral 8x7B leads

Mixtral 8x7B: 32.8 (#269), Phi 3 Mini 128k Instruct: 28.8 (#312)

Coding benchmarks
BenchmarkMixtral 8x7BPhi 3 Mini 128k Instruct
LMArena Coding11261039
BigCodeBench Instruct—29.6%
BigCodeBench Complete—40.6%
HumanEval+39.6%—
MBPP+49.7%—

Reasoning Phi 3 Mini 128k Instruct leads

Mixtral 8x7B: 18.2 (#285), Phi 3 Mini 128k Instruct: 19.6 (#256)

Reasoning benchmarks
BenchmarkMixtral 8x7BPhi 3 Mini 128k Instruct
LMArena Hard Prompts11151028
DTBench49.6%—
Adversarial NLI55.2%—
Epoch Capabilities Index118.47—
ForecastBench56.3—
HellaSwag86.7%—
PIQA83.6%—
WinoGrande77.2%—

Math Phi 3 Mini 128k Instruct leads

Mixtral 8x7B: 18.8 (#289), Phi 3 Mini 128k Instruct: 31.6 (#222)

Math benchmarks
BenchmarkMixtral 8x7BPhi 3 Mini 128k Instruct
LMArena Math11471089
Omni-MATH10.5%—
MATH Level 510%—
GSM8K74.4%—

Knowledge Phi 3 Mini 128k Instruct leads

Mixtral 8x7B: 11.0 (#301), Phi 3 Mini 128k Instruct: 26.8 (#254)

Knowledge benchmarks
BenchmarkMixtral 8x7BPhi 3 Mini 128k Instruct
LMArena Expert1088984
GPQA Diamond30.6%—
MMLU-Pro33.5%—
GPQA (HELM)29.6%—
ARC (AI2) Challenge87.3%—
MMLU70.6%—
OpenBookQA85.8%—
TriviaQA82.2%—

Multilingual Mixtral 8x7B leads

Mixtral 8x7B: 29.6 (#266), Phi 3 Mini 128k Instruct: 25.2 (#285)

Multilingual benchmarks
BenchmarkMixtral 8x7BPhi 3 Mini 128k Instruct
LMArena Non-English10771000
LMArena Chinese10551016
LMArena French11661039
LMArena German11141006
LMArena Japanese931899
LMArena Korean968856
LMArena Russian10901004
LMArena Spanish11111059

Instruction Following Too close to call

Mixtral 8x7B: 51.0 (#297), Phi 3 Mini 128k Instruct: 51.8 (#294)

Instruction Following benchmarks
BenchmarkMixtral 8x7BPhi 3 Mini 128k Instruct
LMArena Instruction Following11091023
IFEval57.5%—

Long Context Mixtral 8x7B leads

Mixtral 8x7B: 33.4 (#260), Phi 3 Mini 128k Instruct: 30.4 (#289)

Long Context benchmarks
BenchmarkMixtral 8x7BPhi 3 Mini 128k Instruct
LMArena Longer Query1103996

Writing & Preference Mixtral 8x7B leads

Mixtral 8x7B: 34.2 (#270), Phi 3 Mini 128k Instruct: 27.1 (#301)

Writing & Preference benchmarks
BenchmarkMixtral 8x7BPhi 3 Mini 128k Instruct
LMArena Text11321050
LMArena Creative Writing11091024
LMArena Multi-Turn1115989
WildBench67.3%—

Frequently asked questions

Is Mixtral 8x7B better than Phi 3 Mini 128k Instruct?

Phi 3 Mini 128k Instruct is the stronger model overall, scoring 29.7 to 27.1 on the Noometry Index.

Is Mixtral 8x7B or Phi 3 Mini 128k Instruct better for coding?

Mixtral 8x7B scores higher on coding benchmarks: 32.8 versus 28.8 in the Noometry coding category.

How many benchmarks do Mixtral 8x7B and Phi 3 Mini 128k Instruct share?

17 benchmarks have published results for both models. Mixtral 8x7B has 38 scored results on Noometry and Phi 3 Mini 128k Instruct has 19.

Related comparisons

Go deeper