Model comparison

Llama 3.2 90B vs Magistral Small

Magistral Small is the stronger model overall, scoring 30.2 to 27.5 on the Noometry Index.

Last verified . 3 shared benchmarks.

Llama 3.2 90B Meta

27.5

Rank #331 Confirmed

Magistral Small Mistral AI

30.2

Rank #296 Confirmed

Summary

  • They share 3 benchmarks with published results for both. Llama 3.2 90B scores higher in 1 category and Magistral Small in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in math, where Magistral Small leads 26.2 to 11.1.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 2.6% for Llama 3.2 90B and 30% for Magistral Small.

Side by side

Llama 3.2 90B and Magistral Small specifications
Llama 3.2 90BMagistral Small
ProviderMetaMistral AI
Noometry Index27.530.2
Released2024-09-242025-06-10
WeightsOpenOpen
Context window—128K
Max output—40K
Input $ / M tokens—$0.50
Output $ / M tokens—$1.50
Results tracked910

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 3.2 90B: —, Magistral Small: 38.4 (#176)

Coding benchmarks
BenchmarkLlama 3.2 90BMagistral Small
SciCode—35.2%

Agentic & Tool Use Not comparable

Llama 3.2 90B: 30.0 (#80), Magistral Small: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 90BMagistral Small
BALROG27.3%—

Reasoning Llama 3.2 90B leads

Llama 3.2 90B: 21.7 (#217), Magistral Small: 6.8 (#350)

Reasoning benchmarks
BenchmarkLlama 3.2 90BMagistral Small
Epoch Capabilities Index125.5133.19
ARC-AGI-2—0%
Kagi LLM Benchmark—6.3%
ARC-AGI-1—5%
CritPt—0.3%
Chess Puzzles—3%
EnigmaEval0.4%—
DTBench—61.3%

Math Magistral Small leads

Llama 3.2 90B: 11.1 (#308), Magistral Small: 26.2 (#261)

Math benchmarks
BenchmarkLlama 3.2 90BMagistral Small
OTIS Mock AIME 2024-20252.6%30%
MATH Level 539.4%—

Knowledge Magistral Small leads

Llama 3.2 90B: 21.7 (#274), Magistral Small: 30.9 (#223)

Knowledge benchmarks
BenchmarkLlama 3.2 90BMagistral Small
GPQA Diamond41%56.1%
MMLU80.3%—

Multimodal Not comparable

Llama 3.2 90B: 25.4 (#124), Magistral Small: —

Multimodal benchmarks
BenchmarkLlama 3.2 90BMagistral Small
LMArena Vision1000—
GeoBench52%—

Frequently asked questions

Is Llama 3.2 90B better than Magistral Small?

Magistral Small is the stronger model overall, scoring 30.2 to 27.5 on the Noometry Index.

How many benchmarks do Llama 3.2 90B and Magistral Small share?

3 benchmarks have published results for both models. Llama 3.2 90B has 9 scored results on Noometry and Magistral Small has 10.

Related comparisons

Go deeper