Model comparison

Llama 3.1 Nemotron Ultra 253b v1 vs Mistral Small 3.1

Llama 3.1 Nemotron Ultra 253b v1 is the stronger model overall, scoring 36.7 to 31.7 on the Noometry Index.

Last verified . 10 shared benchmarks.

Mistral Small 3.1 Mistral AI

31.7

Rank #269 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Llama 3.1 Nemotron Ultra 253b v1 scores higher in 6 categories and Mistral Small 3.1 in 1 category; 5 gaps are clear of the uncertainty.
  • The widest gap is in math, where Llama 3.1 Nemotron Ultra 253b v1 leads 37.5 to 14.7.

Side by side

Llama 3.1 Nemotron Ultra 253b v1 and Mistral Small 3.1 specifications
Llama 3.1 Nemotron Ultra 253b v1Mistral Small 3.1
ProviderNVIDIAMistral AI
Noometry Index36.731.7
Released—2025-03-17
WeightsOpenOpen
Context window—128K
Max output—102K
Input $ / M tokens—$0.35
Output $ / M tokens—$0.56
Results tracked1128

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Llama 3.1 Nemotron Ultra 253b v1: 38.4 (#177), Mistral Small 3.1: 38.3 (#179)

Coding benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Mistral Small 3.1
LMArena Coding13121309

Agentic & Tool Use Not comparable

Llama 3.1 Nemotron Ultra 253b v1: 15.7 (#149), Mistral Small 3.1: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Mistral Small 3.1
Berkeley Function Calling Leaderboard10%—

Reasoning Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 26.3 (#134), Mistral Small 3.1: 19.7 (#254)

Reasoning benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Mistral Small 3.1
LMArena Hard Prompts13161278
Chess Puzzles—1%
Epoch Capabilities Index—127.48

Math Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 37.5 (#152), Mistral Small 3.1: 14.7 (#301)

Math benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Mistral Small 3.1
LMArena Math13601262
OTIS Mock AIME 2024-2025—3.9%
Omni-MATH—24.8%

Knowledge Not comparable

Llama 3.1 Nemotron Ultra 253b v1: —, Mistral Small 3.1: 22.6 (#271)

Knowledge benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Mistral Small 3.1
GPQA Diamond—41.9%
MMLU-Pro—61%
GPQA (HELM)—39.2%
LMArena Expert—1257

Multimodal Not comparable

Llama 3.1 Nemotron Ultra 253b v1: —, Mistral Small 3.1: 33.2 (#99)

Multimodal benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Mistral Small 3.1
LMArena Vision—1136

Multilingual Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 43.1 (#187), Mistral Small 3.1: 41.2 (#209)

Multilingual benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Mistral Small 3.1
LMArena Non-English12821255
LMArena Russian12841263
LMArena Chinese—1253
LMArena French—1273
LMArena German—1266
LMArena Japanese—1208
LMArena Korean—1206
LMArena Spanish—1283

Instruction Following Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 69.0 (#178), Mistral Small 3.1: 63.6 (#230)

Instruction Following benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Mistral Small 3.1
LMArena Instruction Following13081264
IFEval—75%

Long Context Too close to call

Llama 3.1 Nemotron Ultra 253b v1: 39.5 (#177), Mistral Small 3.1: 39.5 (#178)

Long Context benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Mistral Small 3.1
LMArena Longer Query12991299

Writing & Preference Llama 3.1 Nemotron Ultra 253b v1 leads

Llama 3.1 Nemotron Ultra 253b v1: 52.2 (#175), Mistral Small 3.1: 37.0 (#259)

Writing & Preference benchmarks
BenchmarkLlama 3.1 Nemotron Ultra 253b v1Mistral Small 3.1
LMArena Text13201277
LMArena Creative Writing13141253
LMArena Multi-Turn13171270
EQ-Bench Creative Writing—761
WildBench—78.8%

Frequently asked questions

Is Llama 3.1 Nemotron Ultra 253b v1 better than Mistral Small 3.1?

Llama 3.1 Nemotron Ultra 253b v1 is the stronger model overall, scoring 36.7 to 31.7 on the Noometry Index.

Is Llama 3.1 Nemotron Ultra 253b v1 or Mistral Small 3.1 better for coding?

They score almost the same on coding (38.4 vs 38.3); test both on your own repository before choosing.

How many benchmarks do Llama 3.1 Nemotron Ultra 253b v1 and Mistral Small 3.1 share?

10 benchmarks have published results for both models. Llama 3.1 Nemotron Ultra 253b v1 has 11 scored results on Noometry and Mistral Small 3.1 has 28.

Related comparisons

Go deeper