Model comparison

Ministral 8B vs phi-3-medium 14B

phi-3-medium 14B is the stronger model overall, scoring 29.7 to 28.2 on the Noometry Index.

Last verified . 2 shared benchmarks.

Ministral 8B Mistral AI

28.2

Rank #325 Confirmed

phi-3-medium 14B Microsoft

29.7

Rank #306 Reported

Summary

  • They share 2 benchmarks with published results for both. Ministral 8B scores higher in 1 category and phi-3-medium 14B in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Ministral 8B leads 12.6 to 9.1.

Side by side

Ministral 8B and phi-3-medium 14B specifications
Ministral 8Bphi-3-medium 14B
ProviderMistral AIMicrosoft
Noometry Index28.229.7
Released2024-10-012024-04-23
WeightsOpenOpen
Context window262K—
Max output262K—
Input $ / M tokens$0.15—
Output $ / M tokens$0.15—
Results tracked1713

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding phi-3-medium 14B leads

Ministral 8B: 35.0 (#230), phi-3-medium 14B: 36.8 (#201)

Coding benchmarks
BenchmarkMinistral 8Bphi-3-medium 14B
BigCodeBench Instruct—37.6%
LMArena Coding1202—
BigCodeBench Complete—48.7%

Agentic & Tool Use Not comparable

Ministral 8B: 16.4 (#148), phi-3-medium 14B: —

Agentic & Tool Use benchmarks
BenchmarkMinistral 8Bphi-3-medium 14B
Berkeley Function Calling Leaderboard11.1%—

Reasoning Not comparable

Ministral 8B: 18.4 (#281), phi-3-medium 14B: —

Reasoning benchmarks
BenchmarkMinistral 8Bphi-3-medium 14B
LMArena Hard Prompts1191—
DTBench45.7%—
Adversarial NLI—55.8%
BIG-Bench Hard—81.4%
Epoch Capabilities Index—121.23
HellaSwag—82.4%
WinoGrande—81.5%

Math phi-3-medium 14B leads

Ministral 8B: 25.7 (#267), phi-3-medium 14B: 27.3 (#250)

Math benchmarks
BenchmarkMinistral 8Bphi-3-medium 14B
MATH Level 514.9%17.6%
LMArena Math1188—

Knowledge Ministral 8B leads

Ministral 8B: 12.6 (#297), phi-3-medium 14B: 9.1 (#306)

Knowledge benchmarks
BenchmarkMinistral 8Bphi-3-medium 14B
GPQA Diamond27.1%27.6%
Vectara Hallucination Rate7.4%—
LMArena Expert1170—
ARC (AI2) Challenge—91.6%
MMLU—78%
OpenBookQA—87.4%
TriviaQA—73.9%

Multilingual Not comparable

Ministral 8B: 35.1 (#247), phi-3-medium 14B: —

Multilingual benchmarks
BenchmarkMinistral 8Bphi-3-medium 14B
LMArena Non-English1165—
LMArena Chinese1193—
LMArena Russian1195—

Instruction Following Not comparable

Ministral 8B: 60.5 (#250), phi-3-medium 14B: —

Instruction Following benchmarks
BenchmarkMinistral 8Bphi-3-medium 14B
LMArena Instruction Following1161—

Long Context Not comparable

Ministral 8B: 36.7 (#227), phi-3-medium 14B: —

Long Context benchmarks
BenchmarkMinistral 8Bphi-3-medium 14B
LMArena Longer Query1212—

Writing & Preference Not comparable

Ministral 8B: 39.6 (#246), phi-3-medium 14B: —

Writing & Preference benchmarks
BenchmarkMinistral 8Bphi-3-medium 14B
LMArena Text1191—
LMArena Creative Writing1175—
LMArena Multi-Turn1166—

Frequently asked questions

Is Ministral 8B better than phi-3-medium 14B?

phi-3-medium 14B is the stronger model overall, scoring 29.7 to 28.2 on the Noometry Index.

Is Ministral 8B or phi-3-medium 14B better for coding?

phi-3-medium 14B scores higher on coding benchmarks: 36.8 versus 35.0 in the Noometry coding category.

How many benchmarks do Ministral 8B and phi-3-medium 14B share?

2 benchmarks have published results for both models. Ministral 8B has 17 scored results on Noometry and phi-3-medium 14B has 13.

Related comparisons

Go deeper