Model comparison

Llama 3.2 1B vs Magistral Medium

Magistral Medium is the stronger model overall, scoring 35.2 to 20.1 on the Noometry Index. Llama 3.2 1B costs 39× less per token, which makes it the better buy when Magistral Medium's lead doesn't matter for your workload.

Last verified . 13 shared benchmarks.

Llama 3.2 1B Meta

20.1

Rank #354 Confirmed

Magistral Medium Mistral AI

35.2

Rank #227 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Llama 3.2 1B scores higher in 1 category and Magistral Medium in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Magistral Medium leads 33.5 to 7.2.
  • Llama 3.2 1B is cheaper at $0.027 / $0.20 per million input/output tokens, against $2 / $5 for Magistral Medium.
  • Magistral Medium accepts more context: 262K tokens versus 60K.

Side by side

Llama 3.2 1B and Magistral Medium specifications
Llama 3.2 1BMagistral Medium
ProviderMetaMistral AI
Noometry Index20.135.2
Released2024-09-242025-03-17
WeightsOpenOpen
Context window60K262K
Max output54K16K
Input $ / M tokens$0.027$2
Output $ / M tokens$0.20$5
Results tracked2222

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Medium leads

Llama 3.2 1B: 21.1 (#338), Magistral Medium: 39.1 (#161)

Coding benchmarks
BenchmarkLlama 3.2 1BMagistral Medium
LMArena Coding10701319
SciCode—39.2%
BigCodeBench Instruct8.2%—
BigCodeBench Complete11.3%—

Agentic & Tool Use Not comparable

Llama 3.2 1B: 14.6 (#150), Magistral Medium: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 1BMagistral Medium
Berkeley Function Calling Leaderboard10.8%—
BALROG6.6%—

Reasoning Llama 3.2 1B leads

Llama 3.2 1B: 16.2 (#308), Magistral Medium: 8.6 (#348)

Reasoning benchmarks
BenchmarkLlama 3.2 1BMagistral Medium
LMArena Hard Prompts10441267
ARC-AGI-2—0%
Kagi LLM Benchmark—16.2%
ARC-AGI-1—6.1%
CritPt—0.3%
Chess Puzzles0%—
Epoch Capabilities Index101.99—

Math Magistral Medium leads

Llama 3.2 1B: 10.4 (#313), Magistral Medium: 35.1 (#189)

Math benchmarks
BenchmarkLlama 3.2 1BMagistral Medium
LMArena Math10861250
OTIS Mock AIME 2024-20250.6%—

Knowledge Magistral Medium leads

Llama 3.2 1B: 7.2 (#312), Magistral Medium: 33.5 (#202)

Knowledge benchmarks
BenchmarkLlama 3.2 1BMagistral Medium
LMArena Expert10071223
GPQA Diamond23.9%—

Multilingual Magistral Medium leads

Llama 3.2 1B: 23.8 (#292), Magistral Medium: 39.6 (#224)

Multilingual benchmarks
BenchmarkLlama 3.2 1BMagistral Medium
LMArena Non-English9731232
LMArena Chinese9591227
LMArena German10141248
LMArena Russian9411224
LMArena French—1267
LMArena Japanese—1175
LMArena Korean—1125
LMArena Spanish—1271

Instruction Following Magistral Medium leads

Llama 3.2 1B: 52.4 (#290), Magistral Medium: 66.0 (#211)

Instruction Following benchmarks
BenchmarkLlama 3.2 1BMagistral Medium
LMArena Instruction Following10311254

Long Context Magistral Medium leads

Llama 3.2 1B: 31.9 (#274), Magistral Medium: 39.3 (#183)

Long Context benchmarks
BenchmarkLlama 3.2 1BMagistral Medium
LMArena Longer Query10501295

Writing & Preference Magistral Medium leads

Llama 3.2 1B: 21.3 (#310), Magistral Medium: 46.3 (#219)

Writing & Preference benchmarks
BenchmarkLlama 3.2 1BMagistral Medium
LMArena Text10551255
LMArena Creative Writing10331245
LMArena Multi-Turn10301275
EQ-Bench Creative Writing200—

Frequently asked questions

Is Llama 3.2 1B better than Magistral Medium?

Magistral Medium is the stronger model overall, scoring 35.2 to 20.1 on the Noometry Index. Llama 3.2 1B costs 39× less per token, which makes it the better buy when Magistral Medium's lead doesn't matter for your workload.

Which is cheaper, Llama 3.2 1B or Magistral Medium?

Llama 3.2 1B is cheaper. It lists at $0.027 per million input tokens and $0.20 per million output tokens; Magistral Medium lists at $2 and $5.

Is Llama 3.2 1B or Magistral Medium better for coding?

Magistral Medium scores higher on coding benchmarks: 39.1 versus 21.1 in the Noometry coding category.

Which has the bigger context window?

Magistral Medium does, with 262K tokens against 60K.

How many benchmarks do Llama 3.2 1B and Magistral Medium share?

13 benchmarks have published results for both models. Llama 3.2 1B has 22 scored results on Noometry and Magistral Medium has 22.

Related comparisons

Go deeper