Model comparison

Mixtral 8x7B vs Phi 3 Mini 4k Instruct June 2024

Phi 3 Mini 4k Instruct June 2024 is the stronger model overall, scoring 31.3 to 27.1 on the Noometry Index.

Last verified . 15 shared benchmarks.

Mixtral 8x7B Mistral AI

27.1

Rank #334 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Mixtral 8x7B scores higher in 4 categories and Phi 3 Mini 4k Instruct June 2024 in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Phi 3 Mini 4k Instruct June 2024 leads 28.6 to 11.0.

Side by side

Mixtral 8x7B and Phi 3 Mini 4k Instruct June 2024 specifications
Mixtral 8x7BPhi 3 Mini 4k Instruct June 2024
ProviderMistral AIMicrosoft
Noometry Index27.131.3
Released2023-12-11—
WeightsOpenOpen
Context window32K—
Max output32K—
Input $ / M tokens$0.70—
Output $ / M tokens$0.70—
Results tracked3815

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Mixtral 8x7B: 32.8 (#269), Phi 3 Mini 4k Instruct June 2024: 31.8 (#279)

Coding benchmarks
BenchmarkMixtral 8x7BPhi 3 Mini 4k Instruct June 2024
LMArena Coding11261093
HumanEval+39.6%—
MBPP+49.7%—

Reasoning Phi 3 Mini 4k Instruct June 2024 leads

Mixtral 8x7B: 18.2 (#285), Phi 3 Mini 4k Instruct June 2024: 20.8 (#231)

Reasoning benchmarks
BenchmarkMixtral 8x7BPhi 3 Mini 4k Instruct June 2024
LMArena Hard Prompts11151087
DTBench49.6%—
Adversarial NLI55.2%—
Epoch Capabilities Index118.47—
ForecastBench56.3—
HellaSwag86.7%—
PIQA83.6%—
WinoGrande77.2%—

Math Phi 3 Mini 4k Instruct June 2024 leads

Mixtral 8x7B: 18.8 (#289), Phi 3 Mini 4k Instruct June 2024: 33.0 (#208)

Math benchmarks
BenchmarkMixtral 8x7BPhi 3 Mini 4k Instruct June 2024
LMArena Math11471152
Omni-MATH10.5%—
MATH Level 510%—
GSM8K74.4%—

Knowledge Phi 3 Mini 4k Instruct June 2024 leads

Mixtral 8x7B: 11.0 (#301), Phi 3 Mini 4k Instruct June 2024: 28.6 (#244)

Knowledge benchmarks
BenchmarkMixtral 8x7BPhi 3 Mini 4k Instruct June 2024
LMArena Expert10881051
GPQA Diamond30.6%—
MMLU-Pro33.5%—
GPQA (HELM)29.6%—
ARC (AI2) Challenge87.3%—
MMLU70.6%—
OpenBookQA85.8%—
TriviaQA82.2%—

Multilingual Mixtral 8x7B leads

Mixtral 8x7B: 29.6 (#266), Phi 3 Mini 4k Instruct June 2024: 25.9 (#282)

Multilingual benchmarks
BenchmarkMixtral 8x7BPhi 3 Mini 4k Instruct June 2024
LMArena Non-English10771013
LMArena Chinese10551033
LMArena German11141031
LMArena Japanese931954
LMArena Korean968880
LMArena Russian10901019
LMArena French1166—
LMArena Spanish1111—

Instruction Following Phi 3 Mini 4k Instruct June 2024 leads

Mixtral 8x7B: 51.0 (#297), Phi 3 Mini 4k Instruct June 2024: 54.1 (#282)

Instruction Following benchmarks
BenchmarkMixtral 8x7BPhi 3 Mini 4k Instruct June 2024
LMArena Instruction Following11091058
IFEval57.5%—

Long Context Mixtral 8x7B leads

Mixtral 8x7B: 33.4 (#260), Phi 3 Mini 4k Instruct June 2024: 31.7 (#277)

Long Context benchmarks
BenchmarkMixtral 8x7BPhi 3 Mini 4k Instruct June 2024
LMArena Longer Query11031042

Writing & Preference Mixtral 8x7B leads

Mixtral 8x7B: 34.2 (#270), Phi 3 Mini 4k Instruct June 2024: 29.6 (#294)

Writing & Preference benchmarks
BenchmarkMixtral 8x7BPhi 3 Mini 4k Instruct June 2024
LMArena Text11321080
LMArena Creative Writing11091045
LMArena Multi-Turn11151049
WildBench67.3%—

Frequently asked questions

Is Mixtral 8x7B better than Phi 3 Mini 4k Instruct June 2024?

Phi 3 Mini 4k Instruct June 2024 is the stronger model overall, scoring 31.3 to 27.1 on the Noometry Index.

Is Mixtral 8x7B or Phi 3 Mini 4k Instruct June 2024 better for coding?

They score almost the same on coding (32.8 vs 31.8); test both on your own repository before choosing.

How many benchmarks do Mixtral 8x7B and Phi 3 Mini 4k Instruct June 2024 share?

15 benchmarks have published results for both models. Mixtral 8x7B has 38 scored results on Noometry and Phi 3 Mini 4k Instruct June 2024 has 15.

Related comparisons

Go deeper