Model comparison

Llama 3.1 Nemotron 70b Instruct vs Magistral Medium

Llama 3.1 Nemotron 70b Instruct is the stronger model overall, scoring 37.6 to 35.2 on the Noometry Index.

Last verified . 12 shared benchmarks.

Magistral Medium Mistral AI

35.2

Rank #227 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Llama 3.1 Nemotron 70b Instruct scores higher in 5 categories and Magistral Medium in 3 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Llama 3.1 Nemotron 70b Instruct leads 25.0 to 8.6.

Side by side

Llama 3.1 Nemotron 70b Instruct and Magistral Medium specifications
Llama 3.1 Nemotron 70b InstructMagistral Medium
ProviderNVIDIAMistral AI
Noometry Index37.635.2
Released2024-12-182025-03-17
WeightsOpenOpen
Context window—262K
Max output—16K
Input $ / M tokens—$2
Output $ / M tokens—$5
Results tracked1422

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Medium leads

Llama 3.1 Nemotron 70b Instruct: 35.9 (#216), Magistral Medium: 39.1 (#161)

Coding benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructMagistral Medium
LMArena Coding12721319
SciCode—39.2%
BigCodeBench Instruct38.7%—
BigCodeBench Complete48.2%—

Reasoning Llama 3.1 Nemotron 70b Instruct leads

Llama 3.1 Nemotron 70b Instruct: 25.0 (#152), Magistral Medium: 8.6 (#348)

Reasoning benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructMagistral Medium
LMArena Hard Prompts12661267
ARC-AGI-2—0%
Kagi LLM Benchmark—16.2%
ARC-AGI-1—6.1%
CritPt—0.3%

Math Too close to call

Llama 3.1 Nemotron 70b Instruct: 35.5 (#182), Magistral Medium: 35.1 (#189)

Math benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructMagistral Medium
LMArena Math12711250

Knowledge Too close to call

Llama 3.1 Nemotron 70b Instruct: 34.1 (#199), Magistral Medium: 33.5 (#202)

Knowledge benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructMagistral Medium
LMArena Expert12421223

Multilingual Too close to call

Llama 3.1 Nemotron 70b Instruct: 40.5 (#217), Magistral Medium: 39.6 (#224)

Multilingual benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructMagistral Medium
LMArena Non-English12451232
LMArena Chinese12631227
LMArena Russian12271224
LMArena French—1267
LMArena German—1248
LMArena Japanese—1175
LMArena Korean—1125
LMArena Spanish—1271

Instruction Following Too close to call

Llama 3.1 Nemotron 70b Instruct: 65.9 (#213), Magistral Medium: 66.0 (#211)

Instruction Following benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructMagistral Medium
LMArena Instruction Following12521254

Long Context Magistral Medium leads

Llama 3.1 Nemotron 70b Instruct: 37.6 (#215), Magistral Medium: 39.3 (#183)

Long Context benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructMagistral Medium
LMArena Longer Query12381295

Writing & Preference Llama 3.1 Nemotron 70b Instruct leads

Llama 3.1 Nemotron 70b Instruct: 48.4 (#203), Magistral Medium: 46.3 (#219)

Writing & Preference benchmarks
BenchmarkLlama 3.1 Nemotron 70b InstructMagistral Medium
LMArena Text12831255
LMArena Creative Writing12691245
LMArena Multi-Turn12751275

Frequently asked questions

Is Llama 3.1 Nemotron 70b Instruct better than Magistral Medium?

Llama 3.1 Nemotron 70b Instruct is the stronger model overall, scoring 37.6 to 35.2 on the Noometry Index.

Is Llama 3.1 Nemotron 70b Instruct or Magistral Medium better for coding?

Magistral Medium scores higher on coding benchmarks: 39.1 versus 35.9 in the Noometry coding category.

How many benchmarks do Llama 3.1 Nemotron 70b Instruct and Magistral Medium share?

12 benchmarks have published results for both models. Llama 3.1 Nemotron 70b Instruct has 14 scored results on Noometry and Magistral Medium has 22.

Related comparisons

Go deeper