Model comparison

Codestral vs Phi-2

Codestral has enough public results to be ranked (#290); Phi-2 does not yet, so treat this comparison as directional.

Last verified . 2 shared benchmarks.

Codestral Mistral AI

30.6

Rank #290 Reported

Phi-2 Microsoft

—

Unranked

Summary

  • They share 2 benchmarks with published results for both.
  • Phi-2 has downloadable open weights; the other is API-only.

Side by side

Codestral and Phi-2 specifications
CodestralPhi-2
ProviderMistral AIMicrosoft
Noometry Index30.6—
Released2024-05-292023-12-12
WeightsProprietaryOpen
Context window256K—
Max output8K—
Input $ / M tokens$0.30—
Output $ / M tokens$0.90—
Results tracked711

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Codestral: 27.3 (#321), Phi-2: —

Coding benchmarks
BenchmarkCodestralPhi-2
HumanEval+73.8%45.1%
MBPP+61.9%54.2%
Aider Polyglot11.1%—
BigCodeBench Instruct41.8%—
BigCodeBench Complete52.5%—
ALE-Bench137.78—

Reasoning Not comparable

Codestral: 19.8 (#251), Phi-2: —

Reasoning benchmarks
BenchmarkCodestralPhi-2
Kagi LLM Benchmark32.5%—
Adversarial NLI—42.5%
BIG-Bench Hard—59.4%
Epoch Capabilities Index—107.94
HellaSwag—53.6%
WinoGrande—54.7%

Knowledge Not comparable

Codestral: —, Phi-2: —

Knowledge benchmarks
BenchmarkCodestralPhi-2
ARC (AI2) Challenge—75.9%
MMLU—58.4%
OpenBookQA—73.6%
TriviaQA—45.2%

Frequently asked questions

Is Codestral better than Phi-2?

Codestral has enough public results to be ranked (#290); Phi-2 does not yet, so treat this comparison as directional.

How many benchmarks do Codestral and Phi-2 share?

2 benchmarks have published results for both models. Codestral has 7 scored results on Noometry and Phi-2 has 11.

Related comparisons

Go deeper