Model comparison

Mixtral 8x22B vs Phi 3 Mini 4k Instruct June 2024

Phi 3 Mini 4k Instruct June 2024 is the stronger model overall, scoring 31.3 to 27.1 on the Noometry Index.

Last verified . 15 shared benchmarks.

Mixtral 8x22B Mistral AI

27.1

Rank #333 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Mixtral 8x22B scores higher in 4 categories and Phi 3 Mini 4k Instruct June 2024 in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Phi 3 Mini 4k Instruct June 2024 leads 28.6 to 15.1.

Side by side

Mixtral 8x22B and Phi 3 Mini 4k Instruct June 2024 specifications
Mixtral 8x22BPhi 3 Mini 4k Instruct June 2024
ProviderMistral AIMicrosoft
Noometry Index27.131.3
Released2024-04-17—
WeightsOpenOpen
Context window64K—
Max output64K—
Input $ / M tokens$2—
Output $ / M tokens$6—
Results tracked3415

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Phi 3 Mini 4k Instruct June 2024 leads

Mixtral 8x22B: 24.2 (#329), Phi 3 Mini 4k Instruct June 2024: 31.8 (#279)

Coding benchmarks
BenchmarkMixtral 8x22BPhi 3 Mini 4k Instruct June 2024
LMArena Coding11661093
WeirdML3.2%—
BigCodeBench Instruct40.6%—
BigCodeBench Complete50.2%—
HumanEval+72%—
MBPP+64.3%—

Agentic & Tool Use Not comparable

Mixtral 8x22B: 23.1 (#127), Phi 3 Mini 4k Instruct June 2024: —

Agentic & Tool Use benchmarks
BenchmarkMixtral 8x22BPhi 3 Mini 4k Instruct June 2024
Cybench7.5%—

Reasoning Too close to call

Mixtral 8x22B: 19.9 (#248), Phi 3 Mini 4k Instruct June 2024: 20.8 (#231)

Reasoning benchmarks
BenchmarkMixtral 8x22BPhi 3 Mini 4k Instruct June 2024
LMArena Hard Prompts11501087
DTBench55.1%—
Epoch Capabilities Index122.03—
ForecastBench56.3—

Math Phi 3 Mini 4k Instruct June 2024 leads

Mixtral 8x22B: 22.9 (#275), Phi 3 Mini 4k Instruct June 2024: 33.0 (#208)

Math benchmarks
BenchmarkMixtral 8x22BPhi 3 Mini 4k Instruct June 2024
LMArena Math11841152
Omni-MATH16.3%—
MATH Level 524.2%—

Knowledge Phi 3 Mini 4k Instruct June 2024 leads

Mixtral 8x22B: 15.1 (#293), Phi 3 Mini 4k Instruct June 2024: 28.6 (#244)

Knowledge benchmarks
BenchmarkMixtral 8x22BPhi 3 Mini 4k Instruct June 2024
LMArena Expert11131051
GPQA Diamond34.1%—
MMLU-Pro46%—
GPQA (HELM)33.4%—
MMLU77.8%—

Multilingual Mixtral 8x22B leads

Mixtral 8x22B: 32.8 (#255), Phi 3 Mini 4k Instruct June 2024: 25.9 (#282)

Multilingual benchmarks
BenchmarkMixtral 8x22BPhi 3 Mini 4k Instruct June 2024
LMArena Non-English11281013
LMArena Chinese11161033
LMArena German11411031
LMArena Japanese1037954
LMArena Korean1057880
LMArena Russian11581019
LMArena French1166—
LMArena Spanish1151—

Instruction Following Mixtral 8x22B leads

Mixtral 8x22B: 57.7 (#266), Phi 3 Mini 4k Instruct June 2024: 54.1 (#282)

Instruction Following benchmarks
BenchmarkMixtral 8x22BPhi 3 Mini 4k Instruct June 2024
LMArena Instruction Following11471058
IFEval72.4%—

Long Context Mixtral 8x22B leads

Mixtral 8x22B: 34.7 (#247), Phi 3 Mini 4k Instruct June 2024: 31.7 (#277)

Long Context benchmarks
BenchmarkMixtral 8x22BPhi 3 Mini 4k Instruct June 2024
LMArena Longer Query11441042

Writing & Preference Mixtral 8x22B leads

Mixtral 8x22B: 36.9 (#262), Phi 3 Mini 4k Instruct June 2024: 29.6 (#294)

Writing & Preference benchmarks
BenchmarkMixtral 8x22BPhi 3 Mini 4k Instruct June 2024
LMArena Text11621080
LMArena Creative Writing11411045
LMArena Multi-Turn11301049
WildBench71.1%—

Frequently asked questions

Is Mixtral 8x22B better than Phi 3 Mini 4k Instruct June 2024?

Phi 3 Mini 4k Instruct June 2024 is the stronger model overall, scoring 31.3 to 27.1 on the Noometry Index.

Is Mixtral 8x22B or Phi 3 Mini 4k Instruct June 2024 better for coding?

Phi 3 Mini 4k Instruct June 2024 scores higher on coding benchmarks: 31.8 versus 24.2 in the Noometry coding category.

How many benchmarks do Mixtral 8x22B and Phi 3 Mini 4k Instruct June 2024 share?

15 benchmarks have published results for both models. Mixtral 8x22B has 34 scored results on Noometry and Phi 3 Mini 4k Instruct June 2024 has 15.

Related comparisons

Go deeper