Model comparison

Llama 3.2 1B vs Magistral Small

Magistral Small is the stronger model overall, scoring 30.2 to 20.1 on the Noometry Index. Llama 3.2 1B costs 11× less per token, which makes it the better buy when Magistral Small's lead doesn't matter for your workload.

Last verified . 4 shared benchmarks.

Llama 3.2 1B Meta

20.1

Rank #354 Confirmed

Magistral Small Mistral AI

30.2

Rank #296 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Llama 3.2 1B scores higher in 1 category and Magistral Small in 3 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Magistral Small leads 30.9 to 7.2.
  • The biggest single-benchmark swing is GPQA Diamond: 23.9% for Llama 3.2 1B and 56.1% for Magistral Small.
  • Llama 3.2 1B is cheaper at $0.027 / $0.20 per million input/output tokens, against $0.50 / $1.50 for Magistral Small.
  • Magistral Small accepts more context: 128K tokens versus 60K.

Side by side

Llama 3.2 1B and Magistral Small specifications
Llama 3.2 1BMagistral Small
ProviderMetaMistral AI
Noometry Index20.130.2
Released2024-09-242025-06-10
WeightsOpenOpen
Context window60K128K
Max output54K40K
Input $ / M tokens$0.027$0.50
Output $ / M tokens$0.20$1.50
Results tracked2210

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Small leads

Llama 3.2 1B: 21.1 (#338), Magistral Small: 38.4 (#176)

Coding benchmarks
BenchmarkLlama 3.2 1BMagistral Small
SciCode—35.2%
BigCodeBench Instruct8.2%—
LMArena Coding1070—
BigCodeBench Complete11.3%—

Agentic & Tool Use Not comparable

Llama 3.2 1B: 14.6 (#150), Magistral Small: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 1BMagistral Small
Berkeley Function Calling Leaderboard10.8%—
BALROG6.6%—

Reasoning Llama 3.2 1B leads

Llama 3.2 1B: 16.2 (#308), Magistral Small: 6.8 (#350)

Reasoning benchmarks
BenchmarkLlama 3.2 1BMagistral Small
Chess Puzzles0%3%
Epoch Capabilities Index101.99133.19
ARC-AGI-2—0%
Kagi LLM Benchmark—6.3%
ARC-AGI-1—5%
CritPt—0.3%
LMArena Hard Prompts1044—
DTBench—61.3%

Math Magistral Small leads

Llama 3.2 1B: 10.4 (#313), Magistral Small: 26.2 (#261)

Math benchmarks
BenchmarkLlama 3.2 1BMagistral Small
OTIS Mock AIME 2024-20250.6%30%
LMArena Math1086—

Knowledge Magistral Small leads

Llama 3.2 1B: 7.2 (#312), Magistral Small: 30.9 (#223)

Knowledge benchmarks
BenchmarkLlama 3.2 1BMagistral Small
GPQA Diamond23.9%56.1%
LMArena Expert1007—

Multilingual Not comparable

Llama 3.2 1B: 23.8 (#292), Magistral Small: —

Multilingual benchmarks
BenchmarkLlama 3.2 1BMagistral Small
LMArena Non-English973—
LMArena Chinese959—
LMArena German1014—
LMArena Russian941—

Instruction Following Not comparable

Llama 3.2 1B: 52.4 (#290), Magistral Small: —

Instruction Following benchmarks
BenchmarkLlama 3.2 1BMagistral Small
LMArena Instruction Following1031—

Long Context Not comparable

Llama 3.2 1B: 31.9 (#274), Magistral Small: —

Long Context benchmarks
BenchmarkLlama 3.2 1BMagistral Small
LMArena Longer Query1050—

Writing & Preference Not comparable

Llama 3.2 1B: 21.3 (#310), Magistral Small: —

Writing & Preference benchmarks
BenchmarkLlama 3.2 1BMagistral Small
LMArena Text1055—
LMArena Creative Writing1033—
EQ-Bench Creative Writing200—
LMArena Multi-Turn1030—

Frequently asked questions

Is Llama 3.2 1B better than Magistral Small?

Magistral Small is the stronger model overall, scoring 30.2 to 20.1 on the Noometry Index. Llama 3.2 1B costs 11× less per token, which makes it the better buy when Magistral Small's lead doesn't matter for your workload.

Which is cheaper, Llama 3.2 1B or Magistral Small?

Llama 3.2 1B is cheaper. It lists at $0.027 per million input tokens and $0.20 per million output tokens; Magistral Small lists at $0.50 and $1.50.

Is Llama 3.2 1B or Magistral Small better for coding?

Magistral Small scores higher on coding benchmarks: 38.4 versus 21.1 in the Noometry coding category.

Which has the bigger context window?

Magistral Small does, with 128K tokens against 60K.

How many benchmarks do Llama 3.2 1B and Magistral Small share?

4 benchmarks have published results for both models. Llama 3.2 1B has 22 scored results on Noometry and Magistral Small has 10.

Related comparisons

Go deeper