Model comparison

Codestral vs Llama 13b

Codestral is the stronger model overall, scoring 30.6 to 24.4 on the Noometry Index.

Last verified . 0 shared benchmarks.

Codestral Mistral AI

30.6

Rank #290 Reported

Llama 13b Meta

24.4

Rank #348 Confirmed

Summary

  • The widest gap is in coding, where Codestral leads 27.3 to 21.4.
  • Llama 13b has downloadable open weights; the other is API-only.

Side by side

Codestral and Llama 13b specifications
CodestralLlama 13b
ProviderMistral AIMeta
Noometry Index30.624.4
Released2024-05-292023-02-24
WeightsProprietaryOpen
Context window256K—
Max output8K—
Input $ / M tokens$0.30—
Output $ / M tokens$0.90—
Results tracked721

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Codestral leads

Codestral: 27.3 (#321), Llama 13b: 21.4 (#337)

Coding benchmarks
BenchmarkCodestralLlama 13b
Aider Polyglot11.1%—
BigCodeBench Instruct41.8%—
LMArena Coding—683
BigCodeBench Complete52.5%—
ALE-Bench137.78—
HumanEval+73.8%—
MBPP+61.9%—

Reasoning Codestral leads

Codestral: 19.8 (#251), Llama 13b: 14.0 (#329)

Reasoning benchmarks
BenchmarkCodestralLlama 13b
Kagi LLM Benchmark32.5%—
LMArena Hard Prompts—728
BIG-Bench Hard—37.9%
Epoch Capabilities Index—100.58
HellaSwag—79.2%
LAMBADA—75.2%
PIQA—80.1%
WinoGrande—73%

Math Not comparable

Codestral: —, Llama 13b: 26.7 (#256)

Math benchmarks
BenchmarkCodestralLlama 13b
LMArena Math—838
GSM8K—20.6%

Knowledge Not comparable

Codestral: —, Llama 13b: —

Knowledge benchmarks
BenchmarkCodestralLlama 13b
ARC (AI2) Challenge—52.7%
BoolQ—78.7%
MMLU—47.7%
OpenBookQA—56.4%
TriviaQA—77.9%

Multimodal Not comparable

Codestral: —, Llama 13b: —

Multimodal benchmarks
BenchmarkCodestralLlama 13b
ScienceQA—43.3%

Multilingual Not comparable

Codestral: —, Llama 13b: 16.6 (#297)

Multilingual benchmarks
BenchmarkCodestralLlama 13b
LMArena Non-English—819

Instruction Following Not comparable

Codestral: —, Llama 13b: 36.7 (#305)

Instruction Following benchmarks
BenchmarkCodestralLlama 13b
LMArena Instruction Following—781

Writing & Preference Not comparable

Codestral: —, Llama 13b: 13.8 (#312)

Writing & Preference benchmarks
BenchmarkCodestralLlama 13b
LMArena Text—834
LMArena Creative Writing—794
LMArena Multi-Turn—753

Frequently asked questions

Is Codestral better than Llama 13b?

Codestral is the stronger model overall, scoring 30.6 to 24.4 on the Noometry Index.

Is Codestral or Llama 13b better for coding?

Codestral scores higher on coding benchmarks: 27.3 versus 21.4 in the Noometry coding category.

How many benchmarks do Codestral and Llama 13b share?

0 benchmarks have published results for both models. Codestral has 7 scored results on Noometry and Llama 13b has 21.

Related comparisons

Go deeper