Model comparison

Llama 13b vs Magistral Medium

Magistral Medium is the stronger model overall, scoring 35.2 to 24.4 on the Noometry Index.

Last verified . 8 shared benchmarks.

Llama 13b Meta

24.4

Rank #348 Confirmed

Magistral Medium Mistral AI

35.2

Rank #227 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Llama 13b scores higher in 1 category and Magistral Medium in 5 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Magistral Medium leads 46.3 to 13.8.

Side by side

Llama 13b and Magistral Medium specifications
Llama 13bMagistral Medium
ProviderMetaMistral AI
Noometry Index24.435.2
Released2023-02-242025-03-17
WeightsOpenOpen
Context window—262K
Max output—16K
Input $ / M tokens—$2
Output $ / M tokens—$5
Results tracked2122

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Medium leads

Llama 13b: 21.4 (#337), Magistral Medium: 39.1 (#161)

Coding benchmarks
BenchmarkLlama 13bMagistral Medium
LMArena Coding6831319
SciCode—39.2%

Reasoning Llama 13b leads

Llama 13b: 14.0 (#329), Magistral Medium: 8.6 (#348)

Reasoning benchmarks
BenchmarkLlama 13bMagistral Medium
LMArena Hard Prompts7281267
ARC-AGI-2—0%
Kagi LLM Benchmark—16.2%
ARC-AGI-1—6.1%
CritPt—0.3%
BIG-Bench Hard37.9%—
Epoch Capabilities Index100.58—
HellaSwag79.2%—
LAMBADA75.2%—
PIQA80.1%—
WinoGrande73%—

Math Magistral Medium leads

Llama 13b: 26.7 (#256), Magistral Medium: 35.1 (#189)

Math benchmarks
BenchmarkLlama 13bMagistral Medium
LMArena Math8381250
GSM8K20.6%—

Knowledge Not comparable

Llama 13b: —, Magistral Medium: 33.5 (#202)

Knowledge benchmarks
BenchmarkLlama 13bMagistral Medium
LMArena Expert—1223
ARC (AI2) Challenge52.7%—
BoolQ78.7%—
MMLU47.7%—
OpenBookQA56.4%—
TriviaQA77.9%—

Multimodal Not comparable

Llama 13b: —, Magistral Medium: —

Multimodal benchmarks
BenchmarkLlama 13bMagistral Medium
ScienceQA43.3%—

Multilingual Magistral Medium leads

Llama 13b: 16.6 (#297), Magistral Medium: 39.6 (#224)

Multilingual benchmarks
BenchmarkLlama 13bMagistral Medium
LMArena Non-English8191232
LMArena Chinese—1227
LMArena French—1267
LMArena German—1248
LMArena Japanese—1175
LMArena Korean—1125
LMArena Russian—1224
LMArena Spanish—1271

Instruction Following Magistral Medium leads

Llama 13b: 36.7 (#305), Magistral Medium: 66.0 (#211)

Instruction Following benchmarks
BenchmarkLlama 13bMagistral Medium
LMArena Instruction Following7811254

Long Context Not comparable

Llama 13b: —, Magistral Medium: 39.3 (#183)

Long Context benchmarks
BenchmarkLlama 13bMagistral Medium
LMArena Longer Query—1295

Writing & Preference Magistral Medium leads

Llama 13b: 13.8 (#312), Magistral Medium: 46.3 (#219)

Writing & Preference benchmarks
BenchmarkLlama 13bMagistral Medium
LMArena Text8341255
LMArena Creative Writing7941245
LMArena Multi-Turn7531275

Frequently asked questions

Is Llama 13b better than Magistral Medium?

Magistral Medium is the stronger model overall, scoring 35.2 to 24.4 on the Noometry Index.

Is Llama 13b or Magistral Medium better for coding?

Magistral Medium scores higher on coding benchmarks: 39.1 versus 21.4 in the Noometry coding category.

How many benchmarks do Llama 13b and Magistral Medium share?

8 benchmarks have published results for both models. Llama 13b has 21 scored results on Noometry and Magistral Medium has 22.

Related comparisons

Go deeper