Model comparison

Llama 3.2 90B vs Mistral Small 3.2

Mistral Small 3.2 is the stronger model overall, scoring 31.2 to 27.5 on the Noometry Index.

Last verified . 3 shared benchmarks.

Llama 3.2 90B Meta

27.5

Rank #331 Confirmed

Mistral Small 3.2 Mistral AI

31.2

Rank #280 Confirmed

Summary

  • They share 3 benchmarks with published results for both. Llama 3.2 90B scores higher in 1 category and Mistral Small 3.2 in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in math, where Mistral Small 3.2 leads 26.3 to 11.1.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 2.6% for Llama 3.2 90B and 30.3% for Mistral Small 3.2.

Side by side

Llama 3.2 90B and Mistral Small 3.2 specifications
Llama 3.2 90BMistral Small 3.2
ProviderMetaMistral AI
Noometry Index27.531.2
Released2024-09-242025-06-20
WeightsOpenOpen
Context window—256K
Max output—16K
Input $ / M tokens—$0.0938
Output $ / M tokens—$0.25
Results tracked96

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Agentic & Tool Use Not comparable

Llama 3.2 90B: 30.0 (#80), Mistral Small 3.2: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 90BMistral Small 3.2
BALROG27.3%—

Reasoning Llama 3.2 90B leads

Llama 3.2 90B: 21.7 (#217), Mistral Small 3.2: 18.1 (#287)

Reasoning benchmarks
BenchmarkLlama 3.2 90BMistral Small 3.2
Epoch Capabilities Index125.5131.74
Kagi LLM Benchmark—40.4%
Chess Puzzles—1%
EnigmaEval0.4%—

Math Mistral Small 3.2 leads

Llama 3.2 90B: 11.1 (#308), Mistral Small 3.2: 26.3 (#260)

Math benchmarks
BenchmarkLlama 3.2 90BMistral Small 3.2
OTIS Mock AIME 2024-20252.6%30.3%
MATH Level 539.4%—

Knowledge Mistral Small 3.2 leads

Llama 3.2 90B: 21.7 (#274), Mistral Small 3.2: 26.7 (#256)

Knowledge benchmarks
BenchmarkLlama 3.2 90BMistral Small 3.2
GPQA Diamond41%49.1%
MMLU80.3%—

Multimodal Not comparable

Llama 3.2 90B: 25.4 (#124), Mistral Small 3.2: —

Multimodal benchmarks
BenchmarkLlama 3.2 90BMistral Small 3.2
LMArena Vision1000—
GeoBench52%—

Writing & Preference Not comparable

Llama 3.2 90B: —, Mistral Small 3.2: 45.0 (#224)

Writing & Preference benchmarks
BenchmarkLlama 3.2 90BMistral Small 3.2
EQ-Bench Creative Writing—1255

Frequently asked questions

Is Llama 3.2 90B better than Mistral Small 3.2?

Mistral Small 3.2 is the stronger model overall, scoring 31.2 to 27.5 on the Noometry Index.

How many benchmarks do Llama 3.2 90B and Mistral Small 3.2 share?

3 benchmarks have published results for both models. Llama 3.2 90B has 9 scored results on Noometry and Mistral Small 3.2 has 6.

Related comparisons

Go deeper