Model comparison

Codestral vs phi-3-medium 14B

Codestral and phi-3-medium 14B score almost the same on the Noometry Index (30.6 vs 29.7), so choose on price, context window or the category you care about most.

Last verified . 2 shared benchmarks.

Codestral Mistral AI

30.6

Rank #290 Reported

phi-3-medium 14B Microsoft

29.7

Rank #306 Reported

Summary

  • They share 2 benchmarks with published results for both. Codestral scores higher in 0 categories and phi-3-medium 14B in 1 category; one gap is clear of the uncertainty.
  • The widest gap is in coding, where phi-3-medium 14B leads 36.8 to 27.3.
  • phi-3-medium 14B has downloadable open weights; the other is API-only.

Side by side

Codestral and phi-3-medium 14B specifications
Codestralphi-3-medium 14B
ProviderMistral AIMicrosoft
Noometry Index30.629.7
Released2024-05-292024-04-23
WeightsProprietaryOpen
Context window256K—
Max output8K—
Input $ / M tokens$0.30—
Output $ / M tokens$0.90—
Results tracked713

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding phi-3-medium 14B leads

Codestral: 27.3 (#321), phi-3-medium 14B: 36.8 (#201)

Coding benchmarks
BenchmarkCodestralphi-3-medium 14B
BigCodeBench Instruct41.8%37.6%
BigCodeBench Complete52.5%48.7%
Aider Polyglot11.1%—
ALE-Bench137.78—
HumanEval+73.8%—
MBPP+61.9%—

Reasoning Not comparable

Codestral: 19.8 (#251), phi-3-medium 14B: —

Reasoning benchmarks
BenchmarkCodestralphi-3-medium 14B
Kagi LLM Benchmark32.5%—
Adversarial NLI—55.8%
BIG-Bench Hard—81.4%
Epoch Capabilities Index—121.23
HellaSwag—82.4%
WinoGrande—81.5%

Math Not comparable

Codestral: —, phi-3-medium 14B: 27.3 (#250)

Math benchmarks
BenchmarkCodestralphi-3-medium 14B
MATH Level 5—17.6%

Knowledge Not comparable

Codestral: —, phi-3-medium 14B: 9.1 (#306)

Knowledge benchmarks
BenchmarkCodestralphi-3-medium 14B
GPQA Diamond—27.6%
ARC (AI2) Challenge—91.6%
MMLU—78%
OpenBookQA—87.4%
TriviaQA—73.9%

Frequently asked questions

Is Codestral better than phi-3-medium 14B?

Codestral and phi-3-medium 14B score almost the same on the Noometry Index (30.6 vs 29.7), so choose on price, context window or the category you care about most.

Is Codestral or phi-3-medium 14B better for coding?

phi-3-medium 14B scores higher on coding benchmarks: 36.8 versus 27.3 in the Noometry coding category.

How many benchmarks do Codestral and phi-3-medium 14B share?

2 benchmarks have published results for both models. Codestral has 7 scored results on Noometry and phi-3-medium 14B has 13.

Related comparisons

Go deeper