Model comparison

Llama 3-70B vs Magistral Medium

Magistral Medium is the stronger model overall, scoring 35.2 to 28.8 on the Noometry Index.

Last verified . 18 shared benchmarks.

Llama 3-70B Meta

28.8

Rank #323 Confirmed

Magistral Medium Mistral AI

35.2

Rank #227 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Llama 3-70B scores higher in 1 category and Magistral Medium in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Magistral Medium leads 35.1 to 12.8.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 35.1% for Llama 3-70B and 16.2% for Magistral Medium.

Side by side

Llama 3-70B and Magistral Medium specifications
Llama 3-70BMagistral Medium
ProviderMetaMistral AI
Noometry Index28.835.2
Released2024-04-182025-03-17
WeightsOpenOpen
Context window—262K
Max output—16K
Input $ / M tokens—$2
Output $ / M tokens—$5
Results tracked3122

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Medium leads

Llama 3-70B: 35.8 (#218), Magistral Medium: 39.1 (#161)

Coding benchmarks
BenchmarkLlama 3-70BMagistral Medium
LMArena Coding12061319
SciCode—39.2%
BigCodeBench Instruct43.6%—
BigCodeBench Complete54.5%—
HumanEval+72%—
MBPP+69%—

Agentic & Tool Use Not comparable

Llama 3-70B: 21.1 (#139), Magistral Medium: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3-70BMagistral Medium
Cybench5%—

Reasoning Llama 3-70B leads

Llama 3-70B: 18.0 (#288), Magistral Medium: 8.6 (#348)

Reasoning benchmarks
BenchmarkLlama 3-70BMagistral Medium
Kagi LLM Benchmark35.1%16.2%
LMArena Hard Prompts11951267
ARC-AGI-2—0%
ARC-AGI-1—6.1%
CritPt—0.3%
DTBench54.2%—
Epoch Capabilities Index122.93—
ForecastBench57.1—
WinoGrande83.5%—

Math Magistral Medium leads

Llama 3-70B: 12.8 (#305), Magistral Medium: 35.1 (#189)

Math benchmarks
BenchmarkLlama 3-70BMagistral Medium
LMArena Math12181250
OTIS Mock AIME 2024-20254.3%—
MATH Level 522.6%—

Knowledge Magistral Medium leads

Llama 3-70B: 20.8 (#277), Magistral Medium: 33.5 (#202)

Knowledge benchmarks
BenchmarkLlama 3-70BMagistral Medium
LMArena Expert11491223
GPQA Diamond40.6%—
MMLU79.3%—

Multilingual Magistral Medium leads

Llama 3-70B: 33.6 (#251), Magistral Medium: 39.6 (#224)

Multilingual benchmarks
BenchmarkLlama 3-70BMagistral Medium
LMArena Non-English11421232
LMArena Chinese11141227
LMArena French12321267
LMArena German11691248
LMArena Japanese10171175
LMArena Korean10171125
LMArena Russian11591224
LMArena Spanish12411271

Instruction Following Magistral Medium leads

Llama 3-70B: 62.5 (#238), Magistral Medium: 66.0 (#211)

Instruction Following benchmarks
BenchmarkLlama 3-70BMagistral Medium
LMArena Instruction Following11941254

Long Context Magistral Medium leads

Llama 3-70B: 35.6 (#240), Magistral Medium: 39.3 (#183)

Long Context benchmarks
BenchmarkLlama 3-70BMagistral Medium
LMArena Longer Query11741295

Writing & Preference Magistral Medium leads

Llama 3-70B: 42.8 (#231), Magistral Medium: 46.3 (#219)

Writing & Preference benchmarks
BenchmarkLlama 3-70BMagistral Medium
LMArena Text12211255
LMArena Creative Writing12101245
LMArena Multi-Turn12231275

Frequently asked questions

Is Llama 3-70B better than Magistral Medium?

Magistral Medium is the stronger model overall, scoring 35.2 to 28.8 on the Noometry Index.

Is Llama 3-70B or Magistral Medium better for coding?

Magistral Medium scores higher on coding benchmarks: 39.1 versus 35.8 in the Noometry coding category.

How many benchmarks do Llama 3-70B and Magistral Medium share?

18 benchmarks have published results for both models. Llama 3-70B has 31 scored results on Noometry and Magistral Medium has 22.

Related comparisons

Go deeper