Model comparison

Mixtral 8x22B vs Phi 3 Mini 128k Instruct

Phi 3 Mini 128k Instruct is the stronger model overall, scoring 29.7 to 27.1 on the Noometry Index.

Last verified . 19 shared benchmarks.

Mixtral 8x22B Mistral AI

27.1

Rank #333 Confirmed

Phi 3 Mini 128k Instruct Microsoft

29.7

Rank #305 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Mixtral 8x22B scores higher in 5 categories and Phi 3 Mini 128k Instruct in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Phi 3 Mini 128k Instruct leads 26.8 to 15.1.
  • The biggest single-benchmark swing is BigCodeBench Instruct: 40.6% for Mixtral 8x22B and 29.6% for Phi 3 Mini 128k Instruct.

Side by side

Mixtral 8x22B and Phi 3 Mini 128k Instruct specifications
Mixtral 8x22BPhi 3 Mini 128k Instruct
ProviderMistral AIMicrosoft
Noometry Index27.129.7
Released2024-04-172024-04-23
WeightsOpenOpen
Context window64K—
Max output64K—
Input $ / M tokens$2—
Output $ / M tokens$6—
Results tracked3419

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Phi 3 Mini 128k Instruct leads

Mixtral 8x22B: 24.2 (#329), Phi 3 Mini 128k Instruct: 28.8 (#312)

Coding benchmarks
BenchmarkMixtral 8x22BPhi 3 Mini 128k Instruct
BigCodeBench Instruct40.6%29.6%
LMArena Coding11661039
BigCodeBench Complete50.2%40.6%
WeirdML3.2%—
HumanEval+72%—
MBPP+64.3%—

Agentic & Tool Use Not comparable

Mixtral 8x22B: 23.1 (#127), Phi 3 Mini 128k Instruct: —

Agentic & Tool Use benchmarks
BenchmarkMixtral 8x22BPhi 3 Mini 128k Instruct
Cybench7.5%—

Reasoning Too close to call

Mixtral 8x22B: 19.9 (#248), Phi 3 Mini 128k Instruct: 19.6 (#256)

Reasoning benchmarks
BenchmarkMixtral 8x22BPhi 3 Mini 128k Instruct
LMArena Hard Prompts11501028
DTBench55.1%—
Epoch Capabilities Index122.03—
ForecastBench56.3—

Math Phi 3 Mini 128k Instruct leads

Mixtral 8x22B: 22.9 (#275), Phi 3 Mini 128k Instruct: 31.6 (#222)

Math benchmarks
BenchmarkMixtral 8x22BPhi 3 Mini 128k Instruct
LMArena Math11841089
Omni-MATH16.3%—
MATH Level 524.2%—

Knowledge Phi 3 Mini 128k Instruct leads

Mixtral 8x22B: 15.1 (#293), Phi 3 Mini 128k Instruct: 26.8 (#254)

Knowledge benchmarks
BenchmarkMixtral 8x22BPhi 3 Mini 128k Instruct
LMArena Expert1113984
GPQA Diamond34.1%—
MMLU-Pro46%—
GPQA (HELM)33.4%—
MMLU77.8%—

Multilingual Mixtral 8x22B leads

Mixtral 8x22B: 32.8 (#255), Phi 3 Mini 128k Instruct: 25.2 (#285)

Multilingual benchmarks
BenchmarkMixtral 8x22BPhi 3 Mini 128k Instruct
LMArena Non-English11281000
LMArena Chinese11161016
LMArena French11661039
LMArena German11411006
LMArena Japanese1037899
LMArena Korean1057856
LMArena Russian11581004
LMArena Spanish11511059

Instruction Following Mixtral 8x22B leads

Mixtral 8x22B: 57.7 (#266), Phi 3 Mini 128k Instruct: 51.8 (#294)

Instruction Following benchmarks
BenchmarkMixtral 8x22BPhi 3 Mini 128k Instruct
LMArena Instruction Following11471023
IFEval72.4%—

Long Context Mixtral 8x22B leads

Mixtral 8x22B: 34.7 (#247), Phi 3 Mini 128k Instruct: 30.4 (#289)

Long Context benchmarks
BenchmarkMixtral 8x22BPhi 3 Mini 128k Instruct
LMArena Longer Query1144996

Writing & Preference Mixtral 8x22B leads

Mixtral 8x22B: 36.9 (#262), Phi 3 Mini 128k Instruct: 27.1 (#301)

Writing & Preference benchmarks
BenchmarkMixtral 8x22BPhi 3 Mini 128k Instruct
LMArena Text11621050
LMArena Creative Writing11411024
LMArena Multi-Turn1130989
WildBench71.1%—

Frequently asked questions

Is Mixtral 8x22B better than Phi 3 Mini 128k Instruct?

Phi 3 Mini 128k Instruct is the stronger model overall, scoring 29.7 to 27.1 on the Noometry Index.

Is Mixtral 8x22B or Phi 3 Mini 128k Instruct better for coding?

Phi 3 Mini 128k Instruct scores higher on coding benchmarks: 28.8 versus 24.2 in the Noometry coding category.

How many benchmarks do Mixtral 8x22B and Phi 3 Mini 128k Instruct share?

19 benchmarks have published results for both models. Mixtral 8x22B has 34 scored results on Noometry and Phi 3 Mini 128k Instruct has 19.

Related comparisons

Go deeper