Model comparison

Llama 3.2 90B vs Mistral Small

Mistral Small is the stronger model overall, scoring 33.4 to 27.5 on the Noometry Index.

Last verified . 5 shared benchmarks.

Llama 3.2 90B Meta

27.5

Rank #331 Confirmed

Mistral Small Mistral AI

33.4

Rank #243 Confirmed

Summary

  • They share 5 benchmarks with published results for both. Llama 3.2 90B scores higher in 2 categories and Mistral Small in 3 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Mistral Small leads 31.0 to 21.7.
  • The biggest single-benchmark swing is MATH Level 5: 39.4% for Llama 3.2 90B and 46.8% for Mistral Small.

Side by side

Llama 3.2 90B and Mistral Small specifications
Llama 3.2 90BMistral Small
ProviderMetaMistral AI
Noometry Index27.533.4
Released2024-09-242024-02-26
WeightsOpenOpen
Context window—262K
Max output—256K
Input $ / M tokens—$0.15
Output $ / M tokens—$0.60
Results tracked939

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 3.2 90B: —, Mistral Small: 34.0 (#247)

Coding benchmarks
BenchmarkLlama 3.2 90BMistral Small
SciCode—26.5%
BigCodeBench Instruct—36.1%
LiveBench Coding—36.2%
LMArena Coding—1362
BigCodeBench Complete—46.6%
ALE-Bench—497.62

Agentic & Tool Use Llama 3.2 90B leads

Llama 3.2 90B: 30.0 (#80), Mistral Small: 28.1 (#93)

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 90BMistral Small
Berkeley Function Calling Leaderboard—37.1%
BALROG27.3%—

Reasoning Llama 3.2 90B leads

Llama 3.2 90B: 21.7 (#217), Mistral Small: 19.8 (#250)

Reasoning benchmarks
BenchmarkLlama 3.2 90BMistral Small
Kagi LLM Benchmark—37.8%
CritPt—0%
EnigmaEval0.4%—
LiveBench Reasoning—44.8%
LMArena Hard Prompts—1335
DTBench—70.9%
LiveBench Data Analysis—53.7%
LMCA—20.6%
Epoch Capabilities Index125.5—
LiveBench—44%

Math Mistral Small leads

Llama 3.2 90B: 11.1 (#308), Mistral Small: 16.4 (#293)

Math benchmarks
BenchmarkLlama 3.2 90BMistral Small
OTIS Mock AIME 2024-20252.6%5.8%
MATH Level 539.4%46.8%
LiveBench Math—39.9%
LMArena Math—1341

Knowledge Mistral Small leads

Llama 3.2 90B: 21.7 (#274), Mistral Small: 31.0 (#222)

Knowledge benchmarks
BenchmarkLlama 3.2 90BMistral Small
GPQA Diamond41%47.5%
MMLU80.3%68.7%
Vectara Hallucination Rate—5.1%
LMArena Expert—1291

Multimodal Mistral Small leads

Llama 3.2 90B: 25.4 (#124), Mistral Small: 33.5 (#96)

Multimodal benchmarks
BenchmarkLlama 3.2 90BMistral Small
LMArena Vision10001142
GeoBench52%—

Multilingual Not comparable

Llama 3.2 90B: —, Mistral Small: 45.5 (#169)

Multilingual benchmarks
BenchmarkLlama 3.2 90BMistral Small
LMArena Non-English—1315
LMArena Chinese—1340
LMArena French—1337
LMArena German—1340
LMArena Japanese—1275
LMArena Korean—1259
LMArena Russian—1324
LMArena Spanish—1346

Instruction Following Not comparable

Llama 3.2 90B: —, Mistral Small: 66.4 (#209)

Instruction Following benchmarks
BenchmarkLlama 3.2 90BMistral Small
LiveBench Instruction Following—63.7%
LMArena Instruction Following—1310

Long Context Not comparable

Llama 3.2 90B: —, Mistral Small: 40.4 (#156)

Long Context benchmarks
BenchmarkLlama 3.2 90BMistral Small
LMArena Longer Query—1327

Writing & Preference Not comparable

Llama 3.2 90B: —, Mistral Small: 52.5 (#171)

Writing & Preference benchmarks
BenchmarkLlama 3.2 90BMistral Small
LMArena Text—1338
LMArena Creative Writing—1305
LMArena Multi-Turn—1344
LiveBench Language—30.5%

Frequently asked questions

Is Llama 3.2 90B better than Mistral Small?

Mistral Small is the stronger model overall, scoring 33.4 to 27.5 on the Noometry Index.

How many benchmarks do Llama 3.2 90B and Mistral Small share?

5 benchmarks have published results for both models. Llama 3.2 90B has 9 scored results on Noometry and Mistral Small has 39.

Related comparisons

Go deeper