Model comparison

Llama 3.1 Nemotron 51b Instruct vs Mistral Small 3.2

Llama 3.1 Nemotron 51b Instruct is the stronger model overall, scoring 35.9 to 31.2 on the Noometry Index.

Last verified . 0 shared benchmarks.

Mistral Small 3.2 Mistral AI

31.2

Rank #280 Confirmed

Summary

  • The widest gap is in math, where Llama 3.1 Nemotron 51b Instruct leads 34.6 to 26.3.

Side by side

Llama 3.1 Nemotron 51b Instruct and Mistral Small 3.2 specifications
Llama 3.1 Nemotron 51b InstructMistral Small 3.2
ProviderNVIDIAMistral AI
Noometry Index35.931.2
Released—2025-06-20
WeightsOpenOpen
Context window—256K
Max output—16K
Input $ / M tokens—$0.0938
Output $ / M tokens—$0.25
Results tracked126

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 3.1 Nemotron 51b Instruct: 35.6 (#222), Mistral Small 3.2: —

Coding benchmarks
BenchmarkLlama 3.1 Nemotron 51b InstructMistral Small 3.2
LMArena Coding1223—

Reasoning Llama 3.1 Nemotron 51b Instruct leads

Llama 3.1 Nemotron 51b Instruct: 23.5 (#177), Mistral Small 3.2: 18.1 (#287)

Reasoning benchmarks
BenchmarkLlama 3.1 Nemotron 51b InstructMistral Small 3.2
Kagi LLM Benchmark—40.4%
Chess Puzzles—1%
LMArena Hard Prompts1203—
Epoch Capabilities Index—131.74

Math Llama 3.1 Nemotron 51b Instruct leads

Llama 3.1 Nemotron 51b Instruct: 34.6 (#193), Mistral Small 3.2: 26.3 (#260)

Math benchmarks
BenchmarkLlama 3.1 Nemotron 51b InstructMistral Small 3.2
OTIS Mock AIME 2024-2025—30.3%
LMArena Math1230—

Knowledge Llama 3.1 Nemotron 51b Instruct leads

Llama 3.1 Nemotron 51b Instruct: 31.9 (#218), Mistral Small 3.2: 26.7 (#256)

Knowledge benchmarks
BenchmarkLlama 3.1 Nemotron 51b InstructMistral Small 3.2
GPQA Diamond—49.1%
LMArena Expert1167—

Multilingual Not comparable

Llama 3.1 Nemotron 51b Instruct: 36.1 (#241), Mistral Small 3.2: —

Multilingual benchmarks
BenchmarkLlama 3.1 Nemotron 51b InstructMistral Small 3.2
LMArena Non-English1181—
LMArena Chinese1180—
LMArena Russian1187—

Instruction Following Not comparable

Llama 3.1 Nemotron 51b Instruct: 62.9 (#233), Mistral Small 3.2: —

Instruction Following benchmarks
BenchmarkLlama 3.1 Nemotron 51b InstructMistral Small 3.2
LMArena Instruction Following1201—

Long Context Not comparable

Llama 3.1 Nemotron 51b Instruct: 36.5 (#230), Mistral Small 3.2: —

Long Context benchmarks
BenchmarkLlama 3.1 Nemotron 51b InstructMistral Small 3.2
LMArena Longer Query1205—

Writing & Preference Mistral Small 3.2 leads

Llama 3.1 Nemotron 51b Instruct: 43.4 (#229), Mistral Small 3.2: 45.0 (#224)

Writing & Preference benchmarks
BenchmarkLlama 3.1 Nemotron 51b InstructMistral Small 3.2
LMArena Text1228—
LMArena Creative Writing1213—
EQ-Bench Creative Writing—1255
LMArena Multi-Turn1227—

Frequently asked questions

Is Llama 3.1 Nemotron 51b Instruct better than Mistral Small 3.2?

Llama 3.1 Nemotron 51b Instruct is the stronger model overall, scoring 35.9 to 31.2 on the Noometry Index.

How many benchmarks do Llama 3.1 Nemotron 51b Instruct and Mistral Small 3.2 share?

0 benchmarks have published results for both models. Llama 3.1 Nemotron 51b Instruct has 12 scored results on Noometry and Mistral Small 3.2 has 6.

Related comparisons

Go deeper