Model comparison

Llama 3.2 90B vs Mistral Nemo

Llama 3.2 90B is the stronger model overall, scoring 27.5 to 26.4 on the Noometry Index.

Last verified . 4 shared benchmarks.

Llama 3.2 90B Meta

27.5

Rank #331 Confirmed

Mistral Nemo Mistral AI

26.4

Rank #337 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Llama 3.2 90B scores higher in 3 categories and Mistral Nemo in 1 category; 3 gaps are clear of the uncertainty.
  • The widest gap is in math, where Mistral Nemo leads 25.5 to 11.1.
  • The biggest single-benchmark swing is MATH Level 5: 39.4% for Llama 3.2 90B and 10.8% for Mistral Nemo.

Side by side

Llama 3.2 90B and Mistral Nemo specifications
Llama 3.2 90BMistral Nemo
ProviderMetaMistral AI
Noometry Index27.526.4
Released2024-09-242024-07-01
WeightsOpenOpen
Context window—128K
Max output—128K
Input $ / M tokens—$0.15
Output $ / M tokens—$0.15
Results tracked910

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Agentic & Tool Use Llama 3.2 90B leads

Llama 3.2 90B: 30.0 (#80), Mistral Nemo: 23.5 (#125)

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 90BMistral Nemo
BALROG27.3%17.6%
Berkeley Function Calling Leaderboard—27.6%

Reasoning Too close to call

Llama 3.2 90B: 21.7 (#217), Mistral Nemo: 20.7 (#232)

Reasoning benchmarks
BenchmarkLlama 3.2 90BMistral Nemo
Epoch Capabilities Index125.5118.68
EnigmaEval0.4%—
DTBench—48.6%
PIQA—83.5%

Math Mistral Nemo leads

Llama 3.2 90B: 11.1 (#308), Mistral Nemo: 25.5 (#268)

Math benchmarks
BenchmarkLlama 3.2 90BMistral Nemo
MATH Level 539.4%10.8%
OTIS Mock AIME 2024-20252.6%—
GSM8K—84.2%

Knowledge Llama 3.2 90B leads

Llama 3.2 90B: 21.7 (#274), Mistral Nemo: 12.3 (#298)

Knowledge benchmarks
BenchmarkLlama 3.2 90BMistral Nemo
GPQA Diamond41%29.9%
BoolQ—82.5%
MMLU80.3%—

Multimodal Not comparable

Llama 3.2 90B: 25.4 (#124), Mistral Nemo: —

Multimodal benchmarks
BenchmarkLlama 3.2 90BMistral Nemo
LMArena Vision1000—
GeoBench52%—

Writing & Preference Not comparable

Llama 3.2 90B: —, Mistral Nemo: 28.5 (#296)

Writing & Preference benchmarks
BenchmarkLlama 3.2 90BMistral Nemo
EQ-Bench Creative Writing—881

Frequently asked questions

Is Llama 3.2 90B better than Mistral Nemo?

Llama 3.2 90B is the stronger model overall, scoring 27.5 to 26.4 on the Noometry Index.

How many benchmarks do Llama 3.2 90B and Mistral Nemo share?

4 benchmarks have published results for both models. Llama 3.2 90B has 9 scored results on Noometry and Mistral Nemo has 10.

Related comparisons

Go deeper