Model comparison

Llama 13b vs Mistral Medium 3.5

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 24.4 on the Noometry Index.

Last verified . 9 shared benchmarks.

Llama 13b Meta

24.4

Rank #348 Confirmed

Mistral Medium 3.5 Mistral AI

40.2

Rank #152 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Llama 13b scores higher in 0 categories and Mistral Medium 3.5 in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Medium 3.5 leads 58.5 to 13.8.

Side by side

Llama 13b and Mistral Medium 3.5 specifications
Llama 13bMistral Medium 3.5
ProviderMetaMistral AI
Noometry Index24.440.2
Released2023-02-24—
WeightsOpenOpen
Context window—262K
Max output—210K
Input $ / M tokens—$1.50
Output $ / M tokens—$7.50
Results tracked2122

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Medium 3.5 leads

Llama 13b: 21.4 (#337), Mistral Medium 3.5: 36.0 (#213)

Coding benchmarks
BenchmarkLlama 13bMistral Medium 3.5
LMArena Coding6831461
LMArena WebDev—1264

Reasoning Mistral Medium 3.5 leads

Llama 13b: 14.0 (#329), Mistral Medium 3.5: 17.3 (#295)

Reasoning benchmarks
BenchmarkLlama 13bMistral Medium 3.5
LMArena Hard Prompts7281436
Epoch Capabilities Index100.58141.35
Kagi LLM Benchmark—41.4%
NYT Connections (extended)—12.9%
BIG-Bench Hard37.9%—
HellaSwag79.2%—
LAMBADA75.2%—
PIQA80.1%—
WinoGrande73%—

Math Mistral Medium 3.5 leads

Llama 13b: 26.7 (#256), Mistral Medium 3.5: 39.1 (#113)

Math benchmarks
BenchmarkLlama 13bMistral Medium 3.5
LMArena Math8381431
GSM8K20.6%—

Knowledge Not comparable

Llama 13b: —, Mistral Medium 3.5: 40.0 (#126)

Knowledge benchmarks
BenchmarkLlama 13bMistral Medium 3.5
LMArena Expert—1432
ARC (AI2) Challenge52.7%—
BoolQ78.7%—
MMLU47.7%—
OpenBookQA56.4%—
TriviaQA77.9%—

Multimodal Not comparable

Llama 13b: —, Mistral Medium 3.5: 38.3 (#65)

Multimodal benchmarks
BenchmarkLlama 13bMistral Medium 3.5
LMArena Vision—1223
ScienceQA43.3%—

Multilingual Mistral Medium 3.5 leads

Llama 13b: 16.6 (#297), Mistral Medium 3.5: 51.9 (#100)

Multilingual benchmarks
BenchmarkLlama 13bMistral Medium 3.5
LMArena Non-English8191404
LMArena Chinese—1442
LMArena French—1448
LMArena German—1451
LMArena Korean—1385
LMArena Russian—1395
LMArena Spanish—1409

Instruction Following Mistral Medium 3.5 leads

Llama 13b: 36.7 (#305), Mistral Medium 3.5: 74.6 (#90)

Instruction Following benchmarks
BenchmarkLlama 13bMistral Medium 3.5
LMArena Instruction Following7811415

Long Context Not comparable

Llama 13b: —, Mistral Medium 3.5: 43.2 (#103)

Long Context benchmarks
BenchmarkLlama 13bMistral Medium 3.5
LMArena Longer Query—1415

Writing & Preference Mistral Medium 3.5 leads

Llama 13b: 13.8 (#312), Mistral Medium 3.5: 58.5 (#117)

Writing & Preference benchmarks
BenchmarkLlama 13bMistral Medium 3.5
LMArena Text8341421
LMArena Creative Writing7941374
LMArena Multi-Turn7531423
EQ-Bench 4—993

Frequently asked questions

Is Llama 13b better than Mistral Medium 3.5?

Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 24.4 on the Noometry Index.

Is Llama 13b or Mistral Medium 3.5 better for coding?

Mistral Medium 3.5 scores higher on coding benchmarks: 36.0 versus 21.4 in the Noometry coding category.

How many benchmarks do Llama 13b and Mistral Medium 3.5 share?

9 benchmarks have published results for both models. Llama 13b has 21 scored results on Noometry and Mistral Medium 3.5 has 22.

Related comparisons

Go deeper