Model comparison

Mistral vs Phi 3 Small 8k Instruct

Mistral and Phi 3 Small 8k Instruct score almost the same on the Noometry Index (29.9 vs 29.3), so choose on price, context window or the category you care about most.

Last verified . 17 shared benchmarks.

Mistral Mistral AI

29.9

Rank #303 Confirmed

Phi 3 Small 8k Instruct Microsoft

29.3

Rank #314 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Mistral scores higher in 6 categories and Phi 3 Small 8k Instruct in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Phi 3 Small 8k Instruct leads 29.1 to 16.6.
  • Phi 3 Small 8k Instruct has downloadable open weights; the other is API-only.

Side by side

Mistral and Phi 3 Small 8k Instruct specifications
MistralPhi 3 Small 8k Instruct
ProviderMistral AIMicrosoft
Noometry Index29.929.3
Released—2024-04-23
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2232

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral leads

Mistral: 33.8 (#250), Phi 3 Small 8k Instruct: 27.9 (#318)

Coding benchmarks
BenchmarkMistralPhi 3 Small 8k Instruct
LMArena Coding11621101
LiveBench Coding—20.3%

Reasoning Mistral leads

Mistral: 22.2 (#200), Phi 3 Small 8k Instruct: 14.8 (#323)

Reasoning benchmarks
BenchmarkMistralPhi 3 Small 8k Instruct
LMArena Hard Prompts11491100
LiveBench Reasoning—15.9%
LiveBench Data Analysis—30.3%
Adversarial NLI—58.1%
BIG-Bench Hard—79.1%
HellaSwag—77%
LiveBench—24%
WinoGrande—81.5%

Math Phi 3 Small 8k Instruct leads

Mistral: 22.3 (#278), Phi 3 Small 8k Instruct: 27.6 (#248)

Math benchmarks
BenchmarkMistralPhi 3 Small 8k Instruct
LMArena Math11801151
Omni-MATH7.2%—
LiveBench Math—17.6%

Knowledge Phi 3 Small 8k Instruct leads

Mistral: 16.6 (#288), Phi 3 Small 8k Instruct: 29.1 (#240)

Knowledge benchmarks
BenchmarkMistralPhi 3 Small 8k Instruct
LMArena Expert11251067
MMLU-Pro27.7%—
GPQA (HELM)30.3%—
ARC (AI2) Challenge—90.7%
MMLU—75.7%
OpenBookQA—88%
TriviaQA—58.1%

Multilingual Mistral leads

Mistral: 32.8 (#254), Phi 3 Small 8k Instruct: 28.5 (#272)

Multilingual benchmarks
BenchmarkMistralPhi 3 Small 8k Instruct
LMArena Non-English11291058
LMArena Chinese11091061
LMArena French11801135
LMArena German11551080
LMArena Japanese1013966
LMArena Korean1032894
LMArena Russian11681111
LMArena Spanish11431111

Instruction Following Too close to call

Mistral: 52.6 (#288), Phi 3 Small 8k Instruct: 51.9 (#292)

Instruction Following benchmarks
BenchmarkMistralPhi 3 Small 8k Instruct
LMArena Instruction Following11521087
LiveBench Instruction Following—47.2%
IFEval56.8%—

Long Context Mistral leads

Mistral: 35.0 (#245), Phi 3 Small 8k Instruct: 33.0 (#267)

Long Context benchmarks
BenchmarkMistralPhi 3 Small 8k Instruct
LMArena Longer Query11531088

Writing & Preference Mistral leads

Mistral: 37.0 (#260), Phi 3 Small 8k Instruct: 31.1 (#284)

Writing & Preference benchmarks
BenchmarkMistralPhi 3 Small 8k Instruct
LMArena Text11651110
LMArena Creative Writing11581083
LMArena Multi-Turn11471068
WildBench66%—
LiveBench Language—12.9%

Frequently asked questions

Is Mistral better than Phi 3 Small 8k Instruct?

Mistral and Phi 3 Small 8k Instruct score almost the same on the Noometry Index (29.9 vs 29.3), so choose on price, context window or the category you care about most.

Is Mistral or Phi 3 Small 8k Instruct better for coding?

Mistral scores higher on coding benchmarks: 33.8 versus 27.9 in the Noometry coding category.

How many benchmarks do Mistral and Phi 3 Small 8k Instruct share?

17 benchmarks have published results for both models. Mistral has 22 scored results on Noometry and Phi 3 Small 8k Instruct has 32.

Related comparisons

Go deeper