Model comparison

Mistral vs Mistral Small 3.1

Mistral Small 3.1 is the stronger model overall, scoring 31.7 to 29.9 on the Noometry Index.

Last verified . 22 shared benchmarks.

Mistral Mistral AI

29.9

Rank #303 Confirmed

Mistral Small 3.1 Mistral AI

31.7

Rank #269 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Mistral scores higher in 2 categories and Mistral Small 3.1 in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Mistral Small 3.1 leads 63.6 to 52.6.
  • The biggest single-benchmark swing is MMLU-Pro: 27.7% for Mistral and 61% for Mistral Small 3.1.
  • Mistral Small 3.1 has downloadable open weights; the other is API-only.

Side by side

Mistral and Mistral Small 3.1 specifications
MistralMistral Small 3.1
ProviderMistral AIMistral AI
Noometry Index29.931.7
Released—2025-03-17
WeightsProprietaryOpen
Context window—128K
Max output—102K
Input $ / M tokens—$0.35
Output $ / M tokens—$0.56
Results tracked2228

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Small 3.1 leads

Mistral: 33.8 (#250), Mistral Small 3.1: 38.3 (#179)

Coding benchmarks
BenchmarkMistralMistral Small 3.1
LMArena Coding11621309

Reasoning Mistral leads

Mistral: 22.2 (#200), Mistral Small 3.1: 19.7 (#254)

Reasoning benchmarks
BenchmarkMistralMistral Small 3.1
LMArena Hard Prompts11491278
Chess Puzzles—1%
Epoch Capabilities Index—127.48

Math Mistral leads

Mistral: 22.3 (#278), Mistral Small 3.1: 14.7 (#301)

Math benchmarks
BenchmarkMistralMistral Small 3.1
Omni-MATH7.2%24.8%
LMArena Math11801262
OTIS Mock AIME 2024-2025—3.9%

Knowledge Mistral Small 3.1 leads

Mistral: 16.6 (#288), Mistral Small 3.1: 22.6 (#271)

Knowledge benchmarks
BenchmarkMistralMistral Small 3.1
MMLU-Pro27.7%61%
GPQA (HELM)30.3%39.2%
LMArena Expert11251257
GPQA Diamond—41.9%

Multimodal Not comparable

Mistral: —, Mistral Small 3.1: 33.2 (#99)

Multimodal benchmarks
BenchmarkMistralMistral Small 3.1
LMArena Vision—1136

Multilingual Mistral Small 3.1 leads

Mistral: 32.8 (#254), Mistral Small 3.1: 41.2 (#209)

Multilingual benchmarks
BenchmarkMistralMistral Small 3.1
LMArena Non-English11291255
LMArena Chinese11091253
LMArena French11801273
LMArena German11551266
LMArena Japanese10131208
LMArena Korean10321206
LMArena Russian11681263
LMArena Spanish11431283

Instruction Following Mistral Small 3.1 leads

Mistral: 52.6 (#288), Mistral Small 3.1: 63.6 (#230)

Instruction Following benchmarks
BenchmarkMistralMistral Small 3.1
IFEval56.8%75%
LMArena Instruction Following11521264

Long Context Mistral Small 3.1 leads

Mistral: 35.0 (#245), Mistral Small 3.1: 39.5 (#178)

Long Context benchmarks
BenchmarkMistralMistral Small 3.1
LMArena Longer Query11531299

Writing & Preference Too close to call

Mistral: 37.0 (#260), Mistral Small 3.1: 37.0 (#259)

Writing & Preference benchmarks
BenchmarkMistralMistral Small 3.1
LMArena Text11651277
LMArena Creative Writing11581253
WildBench66%78.8%
LMArena Multi-Turn11471270
EQ-Bench Creative Writing—761

Frequently asked questions

Is Mistral better than Mistral Small 3.1?

Mistral Small 3.1 is the stronger model overall, scoring 31.7 to 29.9 on the Noometry Index.

Is Mistral or Mistral Small 3.1 better for coding?

Mistral Small 3.1 scores higher on coding benchmarks: 38.3 versus 33.8 in the Noometry coding category.

How many benchmarks do Mistral and Mistral Small 3.1 share?

22 benchmarks have published results for both models. Mistral has 22 scored results on Noometry and Mistral Small 3.1 has 28.

Related comparisons

Go deeper