Model comparison

Llama-3.3-70B-Instruct vs Magistral Medium

Magistral Medium is the stronger model overall, scoring 35.2 to 30.6 on the Noometry Index. Llama-3.3-70B-Instruct costs 18× less per token, which makes it the better buy when Magistral Medium's lead doesn't matter for your workload.

Last verified . 19 shared benchmarks.

Llama-3.3-70B-Instruct Meta

30.6

Rank #291 Confirmed

Magistral Medium Mistral AI

35.2

Rank #227 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Llama-3.3-70B-Instruct scores higher in 4 categories and Magistral Medium in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Magistral Medium leads 35.1 to 15.3.
  • The biggest single-benchmark swing is SciCode: 26% for Llama-3.3-70B-Instruct and 39.2% for Magistral Medium.
  • Llama-3.3-70B-Instruct is cheaper at $0.10 / $0.32 per million input/output tokens, against $2 / $5 for Magistral Medium.
  • Magistral Medium accepts more context: 262K tokens versus 128K.

Side by side

Llama-3.3-70B-Instruct and Magistral Medium specifications
Llama-3.3-70B-InstructMagistral Medium
ProviderMetaMistral AI
Noometry Index30.635.2
Released2024-12-062025-03-17
WeightsOpenOpen
Context window128K262K
Max output4K16K
Input $ / M tokens$0.10$2
Output $ / M tokens$0.32$5
Results tracked4322

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Medium leads

Llama-3.3-70B-Instruct: 31.0 (#290), Magistral Medium: 39.1 (#161)

Coding benchmarks
BenchmarkLlama-3.3-70B-InstructMagistral Medium
SciCode26%39.2%
LMArena Coding12681319
WeirdML14.4%—
BigCodeBench Instruct46.9%—
LiveBench Coding36.6%—
BigCodeBench Complete57.5%—

Agentic & Tool Use Not comparable

Llama-3.3-70B-Instruct: 25.8 (#105), Magistral Medium: —

Agentic & Tool Use benchmarks
BenchmarkLlama-3.3-70B-InstructMagistral Medium
Berkeley Function Calling Leaderboard31.9%—
BALROG23%—

Reasoning Llama-3.3-70B-Instruct leads

Llama-3.3-70B-Instruct: 14.1 (#327), Magistral Medium: 8.6 (#348)

Reasoning benchmarks
BenchmarkLlama-3.3-70B-InstructMagistral Medium
CritPt0%0.3%
LMArena Hard Prompts12571267
ARC-AGI-2—0%
SimpleBench19.9%—
Kagi LLM Benchmark—16.2%
ARC-AGI-1—6.1%
LiveBench Reasoning50.8%—
DTBench59.5%—
LiveBench Data Analysis49.5%—
LMCA17.5%—
Epoch Capabilities Index127.33—
ForecastBench58.6—
LiveBench50.2%—

Math Magistral Medium leads

Llama-3.3-70B-Instruct: 15.3 (#298), Magistral Medium: 35.1 (#189)

Math benchmarks
BenchmarkLlama-3.3-70B-InstructMagistral Medium
LMArena Math12671250
OTIS Mock AIME 2024-20255.1%—
LiveBench Math42.2%—
MATH Level 541.6%—

Knowledge Magistral Medium leads

Llama-3.3-70B-Instruct: 30.6 (#226), Magistral Medium: 33.5 (#202)

Knowledge benchmarks
BenchmarkLlama-3.3-70B-InstructMagistral Medium
LMArena Expert12251223
GPQA Diamond47.4%—
Confabulations22.8%—
Vectara Hallucination Rate4.1%—
MMLU86.3%—

Multilingual Too close to call

Llama-3.3-70B-Instruct: 39.9 (#220), Magistral Medium: 39.6 (#224)

Multilingual benchmarks
BenchmarkLlama-3.3-70B-InstructMagistral Medium
LMArena Non-English12361232
LMArena Chinese12171227
LMArena French12811267
LMArena German12511248
LMArena Japanese11501175
LMArena Korean11431125
LMArena Russian12521224
LMArena Spanish12701271

Instruction Following Llama-3.3-70B-Instruct leads

Llama-3.3-70B-Instruct: 71.1 (#157), Magistral Medium: 66.0 (#211)

Instruction Following benchmarks
BenchmarkLlama-3.3-70B-InstructMagistral Medium
LMArena Instruction Following12421254
LiveBench Instruction Following82.7%—

Long Context Magistral Medium leads

Llama-3.3-70B-Instruct: 26.4 (#295), Magistral Medium: 39.3 (#183)

Long Context benchmarks
BenchmarkLlama-3.3-70B-InstructMagistral Medium
LMArena Longer Query12561295
Fiction.LiveBench33.3%—

Writing & Preference Llama-3.3-70B-Instruct leads

Llama-3.3-70B-Instruct: 47.6 (#207), Magistral Medium: 46.3 (#219)

Writing & Preference benchmarks
BenchmarkLlama-3.3-70B-InstructMagistral Medium
LMArena Text12741255
LMArena Creative Writing12501245
LMArena Multi-Turn12801275
LiveBench Language39.2%—

Frequently asked questions

Is Llama-3.3-70B-Instruct better than Magistral Medium?

Magistral Medium is the stronger model overall, scoring 35.2 to 30.6 on the Noometry Index. Llama-3.3-70B-Instruct costs 18× less per token, which makes it the better buy when Magistral Medium's lead doesn't matter for your workload.

Which is cheaper, Llama-3.3-70B-Instruct or Magistral Medium?

Llama-3.3-70B-Instruct is cheaper. It lists at $0.10 per million input tokens and $0.32 per million output tokens; Magistral Medium lists at $2 and $5.

Is Llama-3.3-70B-Instruct or Magistral Medium better for coding?

Magistral Medium scores higher on coding benchmarks: 39.1 versus 31.0 in the Noometry coding category.

Which has the bigger context window?

Magistral Medium does, with 262K tokens against 128K.

How many benchmarks do Llama-3.3-70B-Instruct and Magistral Medium share?

19 benchmarks have published results for both models. Llama-3.3-70B-Instruct has 43 scored results on Noometry and Magistral Medium has 22.

Related comparisons

Go deeper