Model comparison

Llama 13b vs Mistral

Mistral is the stronger model overall, scoring 29.9 to 24.4 on the Noometry Index.

Last verified . 8 shared benchmarks.

Llama 13b Meta

24.4

Rank #348 Confirmed

Mistral Mistral AI

29.9

Rank #303 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Llama 13b scores higher in 1 category and Mistral in 5 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral leads 37.0 to 13.8.
  • Llama 13b has downloadable open weights; the other is API-only.

Side by side

Llama 13b and Mistral specifications
Llama 13bMistral
ProviderMetaMistral AI
Noometry Index24.429.9
Released2023-02-24—
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2122

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral leads

Llama 13b: 21.4 (#337), Mistral: 33.8 (#250)

Coding benchmarks
BenchmarkLlama 13bMistral
LMArena Coding6831162

Reasoning Mistral leads

Llama 13b: 14.0 (#329), Mistral: 22.2 (#200)

Reasoning benchmarks
BenchmarkLlama 13bMistral
LMArena Hard Prompts7281149
BIG-Bench Hard37.9%—
Epoch Capabilities Index100.58—
HellaSwag79.2%—
LAMBADA75.2%—
PIQA80.1%—
WinoGrande73%—

Math Llama 13b leads

Llama 13b: 26.7 (#256), Mistral: 22.3 (#278)

Math benchmarks
BenchmarkLlama 13bMistral
LMArena Math8381180
Omni-MATH—7.2%
GSM8K20.6%—

Knowledge Not comparable

Llama 13b: —, Mistral: 16.6 (#288)

Knowledge benchmarks
BenchmarkLlama 13bMistral
MMLU-Pro—27.7%
GPQA (HELM)—30.3%
LMArena Expert—1125
ARC (AI2) Challenge52.7%—
BoolQ78.7%—
MMLU47.7%—
OpenBookQA56.4%—
TriviaQA77.9%—

Multimodal Not comparable

Llama 13b: —, Mistral: —

Multimodal benchmarks
BenchmarkLlama 13bMistral
ScienceQA43.3%—

Multilingual Mistral leads

Llama 13b: 16.6 (#297), Mistral: 32.8 (#254)

Multilingual benchmarks
BenchmarkLlama 13bMistral
LMArena Non-English8191129
LMArena Chinese—1109
LMArena French—1180
LMArena German—1155
LMArena Japanese—1013
LMArena Korean—1032
LMArena Russian—1168
LMArena Spanish—1143

Instruction Following Mistral leads

Llama 13b: 36.7 (#305), Mistral: 52.6 (#288)

Instruction Following benchmarks
BenchmarkLlama 13bMistral
LMArena Instruction Following7811152
IFEval—56.8%

Long Context Not comparable

Llama 13b: —, Mistral: 35.0 (#245)

Long Context benchmarks
BenchmarkLlama 13bMistral
LMArena Longer Query—1153

Writing & Preference Mistral leads

Llama 13b: 13.8 (#312), Mistral: 37.0 (#260)

Writing & Preference benchmarks
BenchmarkLlama 13bMistral
LMArena Text8341165
LMArena Creative Writing7941158
LMArena Multi-Turn7531147
WildBench—66%

Frequently asked questions

Is Llama 13b better than Mistral?

Mistral is the stronger model overall, scoring 29.9 to 24.4 on the Noometry Index.

Is Llama 13b or Mistral better for coding?

Mistral scores higher on coding benchmarks: 33.8 versus 21.4 in the Noometry coding category.

How many benchmarks do Llama 13b and Mistral share?

8 benchmarks have published results for both models. Llama 13b has 21 scored results on Noometry and Mistral has 22.

Related comparisons

Go deeper