Model comparison

Llama2 70b Steerlm Chat vs Magistral Medium

Magistral Medium is the stronger model overall, scoring 35.2 to 31.8 on the Noometry Index.

Last verified . 9 shared benchmarks.

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Magistral Medium Mistral AI

35.2

Rank #227 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Llama2 70b Steerlm Chat scores higher in 1 category and Magistral Medium in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Magistral Medium leads 46.3 to 31.6.

Side by side

Llama2 70b Steerlm Chat and Magistral Medium specifications
Llama2 70b Steerlm ChatMagistral Medium
ProviderNVIDIAMistral AI
Noometry Index31.835.2
Released—2025-03-17
WeightsOpenOpen
Context window—262K
Max output—16K
Input $ / M tokens—$2
Output $ / M tokens—$5
Results tracked922

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Medium leads

Llama2 70b Steerlm Chat: 29.9 (#300), Magistral Medium: 39.1 (#161)

Coding benchmarks
BenchmarkLlama2 70b Steerlm ChatMagistral Medium
LMArena Coding10251319
SciCode—39.2%

Reasoning Llama2 70b Steerlm Chat leads

Llama2 70b Steerlm Chat: 20.0 (#246), Magistral Medium: 8.6 (#348)

Reasoning benchmarks
BenchmarkLlama2 70b Steerlm ChatMagistral Medium
LMArena Hard Prompts10471267
ARC-AGI-2—0%
Kagi LLM Benchmark—16.2%
ARC-AGI-1—6.1%
CritPt—0.3%

Math Magistral Medium leads

Llama2 70b Steerlm Chat: 31.3 (#226), Magistral Medium: 35.1 (#189)

Math benchmarks
BenchmarkLlama2 70b Steerlm ChatMagistral Medium
LMArena Math10721250

Knowledge Not comparable

Llama2 70b Steerlm Chat: —, Magistral Medium: 33.5 (#202)

Knowledge benchmarks
BenchmarkLlama2 70b Steerlm ChatMagistral Medium
LMArena Expert—1223

Multilingual Magistral Medium leads

Llama2 70b Steerlm Chat: 28.8 (#270), Magistral Medium: 39.6 (#224)

Multilingual benchmarks
BenchmarkLlama2 70b Steerlm ChatMagistral Medium
LMArena Non-English10631232
LMArena Chinese—1227
LMArena French—1267
LMArena German—1248
LMArena Japanese—1175
LMArena Korean—1125
LMArena Russian—1224
LMArena Spanish—1271

Instruction Following Magistral Medium leads

Llama2 70b Steerlm Chat: 54.2 (#279), Magistral Medium: 66.0 (#211)

Instruction Following benchmarks
BenchmarkLlama2 70b Steerlm ChatMagistral Medium
LMArena Instruction Following10601254

Long Context Magistral Medium leads

Llama2 70b Steerlm Chat: 30.4 (#288), Magistral Medium: 39.3 (#183)

Long Context benchmarks
BenchmarkLlama2 70b Steerlm ChatMagistral Medium
LMArena Longer Query9981295

Writing & Preference Magistral Medium leads

Llama2 70b Steerlm Chat: 31.6 (#283), Magistral Medium: 46.3 (#219)

Writing & Preference benchmarks
BenchmarkLlama2 70b Steerlm ChatMagistral Medium
LMArena Text10981255
LMArena Creative Writing10911245
LMArena Multi-Turn10581275

Frequently asked questions

Is Llama2 70b Steerlm Chat better than Magistral Medium?

Magistral Medium is the stronger model overall, scoring 35.2 to 31.8 on the Noometry Index.

Is Llama2 70b Steerlm Chat or Magistral Medium better for coding?

Magistral Medium scores higher on coding benchmarks: 39.1 versus 29.9 in the Noometry coding category.

How many benchmarks do Llama2 70b Steerlm Chat and Magistral Medium share?

9 benchmarks have published results for both models. Llama2 70b Steerlm Chat has 9 scored results on Noometry and Magistral Medium has 22.

Related comparisons

Go deeper