Model comparison

Llama 13b vs Mistral Medium 3.1

Mistral Medium 3.1 is the stronger model overall, scoring 31.9 to 24.4 on the Noometry Index.

Last verified . 0 shared benchmarks.

Llama 13b Meta

24.4

Rank #348 Confirmed

Mistral Medium 3.1 Mistral AI

31.9

Rank #266 Reported

Summary

  • The widest gap is in writing & preference, where Mistral Medium 3.1 leads 55.5 to 13.8.
  • Llama 13b has downloadable open weights; the other is API-only.

Side by side

Llama 13b and Mistral Medium 3.1 specifications
Llama 13bMistral Medium 3.1
ProviderMetaMistral AI
Noometry Index24.431.9
Released2023-02-24—
WeightsOpenProprietary
Context window—131K
Max output—105K
Input $ / M tokens—$0.40
Output $ / M tokens—$2
Results tracked213

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 13b: 21.4 (#337), Mistral Medium 3.1: —

Coding benchmarks
BenchmarkLlama 13bMistral Medium 3.1
LMArena Coding683—

Reasoning Llama 13b leads

Llama 13b: 14.0 (#329), Mistral Medium 3.1: 10.6 (#341)

Reasoning benchmarks
BenchmarkLlama 13bMistral Medium 3.1
NYT Connections (extended)—6.5%
Thematic Generalization—20.3%
LMArena Hard Prompts728—
BIG-Bench Hard37.9%—
Epoch Capabilities Index100.58—
HellaSwag79.2%—
LAMBADA75.2%—
PIQA80.1%—
WinoGrande73%—

Math Not comparable

Llama 13b: 26.7 (#256), Mistral Medium 3.1: —

Math benchmarks
BenchmarkLlama 13bMistral Medium 3.1
LMArena Math838—
GSM8K20.6%—

Knowledge Not comparable

Llama 13b: —, Mistral Medium 3.1: —

Knowledge benchmarks
BenchmarkLlama 13bMistral Medium 3.1
ARC (AI2) Challenge52.7%—
BoolQ78.7%—
MMLU47.7%—
OpenBookQA56.4%—
TriviaQA77.9%—

Multimodal Not comparable

Llama 13b: —, Mistral Medium 3.1: —

Multimodal benchmarks
BenchmarkLlama 13bMistral Medium 3.1
ScienceQA43.3%—

Multilingual Not comparable

Llama 13b: 16.6 (#297), Mistral Medium 3.1: —

Multilingual benchmarks
BenchmarkLlama 13bMistral Medium 3.1
LMArena Non-English819—

Instruction Following Not comparable

Llama 13b: 36.7 (#305), Mistral Medium 3.1: —

Instruction Following benchmarks
BenchmarkLlama 13bMistral Medium 3.1
LMArena Instruction Following781—

Writing & Preference Mistral Medium 3.1 leads

Llama 13b: 13.8 (#312), Mistral Medium 3.1: 55.5 (#145)

Writing & Preference benchmarks
BenchmarkLlama 13bMistral Medium 3.1
LMArena Text834—
LMArena Creative Writing794—
EQ-Bench Creative Writing—1476
LMArena Multi-Turn753—

Frequently asked questions

Is Llama 13b better than Mistral Medium 3.1?

Mistral Medium 3.1 is the stronger model overall, scoring 31.9 to 24.4 on the Noometry Index.

How many benchmarks do Llama 13b and Mistral Medium 3.1 share?

0 benchmarks have published results for both models. Llama 13b has 21 scored results on Noometry and Mistral Medium 3.1 has 3.

Related comparisons

Go deeper