Model comparison

Mistral Small 3 vs phi-3-medium 14B

Mistral Small 3 is the stronger model overall, scoring 31.2 to 29.7 on the Noometry Index.

Last verified . 4 shared benchmarks.

Mistral Small 3 Mistral AI

31.2

Rank #278 Confirmed

phi-3-medium 14B Microsoft

29.7

Rank #306 Reported

Summary

  • They share 4 benchmarks with published results for both. Mistral Small 3 scores higher in 1 category and phi-3-medium 14B in 2 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Mistral Small 3 leads 25.1 to 9.1.
  • The biggest single-benchmark swing is GPQA Diamond: 47.3% for Mistral Small 3 and 27.6% for phi-3-medium 14B.

Side by side

Mistral Small 3 and phi-3-medium 14B specifications
Mistral Small 3phi-3-medium 14B
ProviderMistral AIMicrosoft
Noometry Index31.229.7
Released2025-01-302024-04-23
WeightsOpenOpen
Context window33K—
Max output16K—
Input $ / M tokens$0.05—
Output $ / M tokens$0.08—
Results tracked2413

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Mistral Small 3: 36.5 (#207), phi-3-medium 14B: 36.8 (#201)

Coding benchmarks
BenchmarkMistral Small 3phi-3-medium 14B
BigCodeBench Instruct45.3%37.6%
BigCodeBench Complete50.4%48.7%
LMArena Coding1246—

Reasoning Not comparable

Mistral Small 3: 18.9 (#273), phi-3-medium 14B: —

Reasoning benchmarks
BenchmarkMistral Small 3phi-3-medium 14B
Epoch Capabilities Index127.07121.23
Chess Puzzles0%—
LMArena Hard Prompts1233—
Adversarial NLI—55.8%
BIG-Bench Hard—81.4%
HellaSwag—82.4%
WinoGrande—81.5%

Math phi-3-medium 14B leads

Mistral Small 3: 16.3 (#295), phi-3-medium 14B: 27.3 (#250)

Math benchmarks
BenchmarkMistral Small 3phi-3-medium 14B
OTIS Mock AIME 2024-20256.7%—
LMArena Math1240—
MATH Level 5—17.6%

Knowledge Mistral Small 3 leads

Mistral Small 3: 25.1 (#263), phi-3-medium 14B: 9.1 (#306)

Knowledge benchmarks
BenchmarkMistral Small 3phi-3-medium 14B
GPQA Diamond47.3%27.6%
Confabulations25.2%—
LMArena Expert1202—
ARC (AI2) Challenge—91.6%
MMLU—78%
OpenBookQA—87.4%
TriviaQA—73.9%

Multilingual Not comparable

Mistral Small 3: 37.3 (#236), phi-3-medium 14B: —

Multilingual benchmarks
BenchmarkMistral Small 3phi-3-medium 14B
LMArena Non-English1198—
LMArena Chinese1204—
LMArena French1203—
LMArena German1211—
LMArena Japanese1111—
LMArena Korean1188—
LMArena Russian1216—

Instruction Following Not comparable

Mistral Small 3: 63.7 (#229), phi-3-medium 14B: —

Instruction Following benchmarks
BenchmarkMistral Small 3phi-3-medium 14B
LMArena Instruction Following1214—

Long Context Not comparable

Mistral Small 3: 37.8 (#211), phi-3-medium 14B: —

Long Context benchmarks
BenchmarkMistral Small 3phi-3-medium 14B
LMArena Longer Query1246—

Writing & Preference Not comparable

Mistral Small 3: 32.2 (#280), phi-3-medium 14B: —

Writing & Preference benchmarks
BenchmarkMistral Small 3phi-3-medium 14B
LMArena Text1234—
LMArena Creative Writing1195—
EQ-Bench Creative Writing707—
LMArena Multi-Turn1217—

Frequently asked questions

Is Mistral Small 3 better than phi-3-medium 14B?

Mistral Small 3 is the stronger model overall, scoring 31.2 to 29.7 on the Noometry Index.

Is Mistral Small 3 or phi-3-medium 14B better for coding?

They score almost the same on coding (36.5 vs 36.8); test both on your own repository before choosing.

How many benchmarks do Mistral Small 3 and phi-3-medium 14B share?

4 benchmarks have published results for both models. Mistral Small 3 has 24 scored results on Noometry and phi-3-medium 14B has 13.

Related comparisons

Go deeper