Model comparison

Mistral Nemo vs Mistral Small 3.1

Mistral Small 3.1 is the stronger model overall, scoring 31.7 to 26.4 on the Noometry Index. Mistral Nemo costs 2.7× less per token, which makes it the better buy when Mistral Small 3.1's lead doesn't matter for your workload.

Last verified . 3 shared benchmarks.

Mistral Nemo Mistral AI

26.4

Rank #337 Confirmed

Mistral Small 3.1 Mistral AI

31.7

Rank #269 Confirmed

Summary

  • They share 3 benchmarks with published results for both. Mistral Nemo scores higher in 2 categories and Mistral Small 3.1 in 2 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in math, where Mistral Nemo leads 25.5 to 14.7.
  • The biggest single-benchmark swing is GPQA Diamond: 29.9% for Mistral Nemo and 41.9% for Mistral Small 3.1.
  • Mistral Nemo is cheaper at $0.15 / $0.15 per million input/output tokens, against $0.35 / $0.56 for Mistral Small 3.1.

Side by side

Mistral Nemo and Mistral Small 3.1 specifications
Mistral NemoMistral Small 3.1
ProviderMistral AIMistral AI
Noometry Index26.431.7
Released2024-07-012025-03-17
WeightsOpenOpen
Context window128K128K
Max output128K102K
Input $ / M tokens$0.15$0.35
Output $ / M tokens$0.15$0.56
Results tracked1028

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Mistral Nemo: —, Mistral Small 3.1: 38.3 (#179)

Coding benchmarks
BenchmarkMistral NemoMistral Small 3.1
LMArena Coding—1309

Agentic & Tool Use Not comparable

Mistral Nemo: 23.5 (#125), Mistral Small 3.1: —

Agentic & Tool Use benchmarks
BenchmarkMistral NemoMistral Small 3.1
Berkeley Function Calling Leaderboard27.6%—
BALROG17.6%—

Reasoning Mistral Nemo leads

Mistral Nemo: 20.7 (#232), Mistral Small 3.1: 19.7 (#254)

Reasoning benchmarks
BenchmarkMistral NemoMistral Small 3.1
Epoch Capabilities Index118.68127.48
Chess Puzzles—1%
LMArena Hard Prompts—1278
DTBench48.6%—
PIQA83.5%—

Math Mistral Nemo leads

Mistral Nemo: 25.5 (#268), Mistral Small 3.1: 14.7 (#301)

Math benchmarks
BenchmarkMistral NemoMistral Small 3.1
OTIS Mock AIME 2024-2025—3.9%
Omni-MATH—24.8%
LMArena Math—1262
MATH Level 510.8%—
GSM8K84.2%—

Knowledge Mistral Small 3.1 leads

Mistral Nemo: 12.3 (#298), Mistral Small 3.1: 22.6 (#271)

Knowledge benchmarks
BenchmarkMistral NemoMistral Small 3.1
GPQA Diamond29.9%41.9%
MMLU-Pro—61%
GPQA (HELM)—39.2%
LMArena Expert—1257
BoolQ82.5%—

Multimodal Not comparable

Mistral Nemo: —, Mistral Small 3.1: 33.2 (#99)

Multimodal benchmarks
BenchmarkMistral NemoMistral Small 3.1
LMArena Vision—1136

Multilingual Not comparable

Mistral Nemo: —, Mistral Small 3.1: 41.2 (#209)

Multilingual benchmarks
BenchmarkMistral NemoMistral Small 3.1
LMArena Non-English—1255
LMArena Chinese—1253
LMArena French—1273
LMArena German—1266
LMArena Japanese—1208
LMArena Korean—1206
LMArena Russian—1263
LMArena Spanish—1283

Instruction Following Not comparable

Mistral Nemo: —, Mistral Small 3.1: 63.6 (#230)

Instruction Following benchmarks
BenchmarkMistral NemoMistral Small 3.1
IFEval—75%
LMArena Instruction Following—1264

Long Context Not comparable

Mistral Nemo: —, Mistral Small 3.1: 39.5 (#178)

Long Context benchmarks
BenchmarkMistral NemoMistral Small 3.1
LMArena Longer Query—1299

Writing & Preference Mistral Small 3.1 leads

Mistral Nemo: 28.5 (#296), Mistral Small 3.1: 37.0 (#259)

Writing & Preference benchmarks
BenchmarkMistral NemoMistral Small 3.1
EQ-Bench Creative Writing881761
LMArena Text—1277
LMArena Creative Writing—1253
WildBench—78.8%
LMArena Multi-Turn—1270

Frequently asked questions

Is Mistral Nemo better than Mistral Small 3.1?

Mistral Small 3.1 is the stronger model overall, scoring 31.7 to 26.4 on the Noometry Index. Mistral Nemo costs 2.7× less per token, which makes it the better buy when Mistral Small 3.1's lead doesn't matter for your workload.

Which is cheaper, Mistral Nemo or Mistral Small 3.1?

Mistral Nemo is cheaper. It lists at $0.15 per million input tokens and $0.15 per million output tokens; Mistral Small 3.1 lists at $0.35 and $0.56.

Which has the bigger context window?

Both accept 128K tokens.

How many benchmarks do Mistral Nemo and Mistral Small 3.1 share?

3 benchmarks have published results for both models. Mistral Nemo has 10 scored results on Noometry and Mistral Small 3.1 has 28.

Related comparisons

Go deeper