Model comparison

Mistral Small vs Mistral Small 3.1

Mistral Small is the stronger model overall, scoring 33.4 to 31.7 on the Noometry Index.

Last verified . 20 shared benchmarks.

Mistral Small Mistral AI

33.4

Rank #243 Confirmed

Mistral Small 3.1 Mistral AI

31.7

Rank #269 Confirmed

Summary

  • They share 20 benchmarks with published results for both. Mistral Small scores higher in 8 categories and Mistral Small 3.1 in 1 category; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Small leads 52.5 to 37.0.
  • The biggest single-benchmark swing is GPQA Diamond: 47.5% for Mistral Small and 41.9% for Mistral Small 3.1.
  • Mistral Small is cheaper at $0.15 / $0.60 per million input/output tokens, against $0.35 / $0.56 for Mistral Small 3.1.
  • Mistral Small accepts more context: 262K tokens versus 128K.

Side by side

Mistral Small and Mistral Small 3.1 specifications
Mistral SmallMistral Small 3.1
ProviderMistral AIMistral AI
Noometry Index33.431.7
Released2024-02-262025-03-17
WeightsOpenOpen
Context window262K128K
Max output256K102K
Input $ / M tokens$0.15$0.35
Output $ / M tokens$0.60$0.56
Results tracked3928

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Small 3.1 leads

Mistral Small: 34.0 (#247), Mistral Small 3.1: 38.3 (#179)

Coding benchmarks
BenchmarkMistral SmallMistral Small 3.1
LMArena Coding13621309
SciCode26.5%—
BigCodeBench Instruct36.1%—
LiveBench Coding36.2%—
BigCodeBench Complete46.6%—
ALE-Bench497.62—

Agentic & Tool Use Not comparable

Mistral Small: 28.1 (#93), Mistral Small 3.1: —

Agentic & Tool Use benchmarks
BenchmarkMistral SmallMistral Small 3.1
Berkeley Function Calling Leaderboard37.1%—

Reasoning Too close to call

Mistral Small: 19.8 (#250), Mistral Small 3.1: 19.7 (#254)

Reasoning benchmarks
BenchmarkMistral SmallMistral Small 3.1
LMArena Hard Prompts13351278
Kagi LLM Benchmark37.8%—
CritPt0%—
Chess Puzzles—1%
LiveBench Reasoning44.8%—
DTBench70.9%—
LiveBench Data Analysis53.7%—
LMCA20.6%—
Epoch Capabilities Index—127.48
LiveBench44%—

Math Mistral Small leads

Mistral Small: 16.4 (#293), Mistral Small 3.1: 14.7 (#301)

Math benchmarks
BenchmarkMistral SmallMistral Small 3.1
OTIS Mock AIME 2024-20255.8%3.9%
LMArena Math13411262
Omni-MATH—24.8%
LiveBench Math39.9%—
MATH Level 546.8%—

Knowledge Mistral Small leads

Mistral Small: 31.0 (#222), Mistral Small 3.1: 22.6 (#271)

Knowledge benchmarks
BenchmarkMistral SmallMistral Small 3.1
GPQA Diamond47.5%41.9%
LMArena Expert12911257
MMLU-Pro—61%
Vectara Hallucination Rate5.1%—
GPQA (HELM)—39.2%
MMLU68.7%—

Multimodal Too close to call

Mistral Small: 33.5 (#96), Mistral Small 3.1: 33.2 (#99)

Multimodal benchmarks
BenchmarkMistral SmallMistral Small 3.1
LMArena Vision11421136

Multilingual Mistral Small leads

Mistral Small: 45.5 (#169), Mistral Small 3.1: 41.2 (#209)

Multilingual benchmarks
BenchmarkMistral SmallMistral Small 3.1
LMArena Non-English13151255
LMArena Chinese13401253
LMArena French13371273
LMArena German13401266
LMArena Japanese12751208
LMArena Korean12591206
LMArena Russian13241263
LMArena Spanish13461283

Instruction Following Mistral Small leads

Mistral Small: 66.4 (#209), Mistral Small 3.1: 63.6 (#230)

Instruction Following benchmarks
BenchmarkMistral SmallMistral Small 3.1
LMArena Instruction Following13101264
LiveBench Instruction Following63.7%—
IFEval—75%

Long Context Too close to call

Mistral Small: 40.4 (#156), Mistral Small 3.1: 39.5 (#178)

Long Context benchmarks
BenchmarkMistral SmallMistral Small 3.1
LMArena Longer Query13271299

Writing & Preference Mistral Small leads

Mistral Small: 52.5 (#171), Mistral Small 3.1: 37.0 (#259)

Writing & Preference benchmarks
BenchmarkMistral SmallMistral Small 3.1
LMArena Text13381277
LMArena Creative Writing13051253
LMArena Multi-Turn13441270
EQ-Bench Creative Writing—761
WildBench—78.8%
LiveBench Language30.5%—

Frequently asked questions

Is Mistral Small better than Mistral Small 3.1?

Mistral Small is the stronger model overall, scoring 33.4 to 31.7 on the Noometry Index.

Which is cheaper, Mistral Small or Mistral Small 3.1?

Mistral Small is cheaper. It lists at $0.15 per million input tokens and $0.60 per million output tokens; Mistral Small 3.1 lists at $0.35 and $0.56.

Is Mistral Small or Mistral Small 3.1 better for coding?

Mistral Small 3.1 scores higher on coding benchmarks: 38.3 versus 34.0 in the Noometry coding category.

Which has the bigger context window?

Mistral Small does, with 262K tokens against 128K.

How many benchmarks do Mistral Small and Mistral Small 3.1 share?

20 benchmarks have published results for both models. Mistral Small has 39 scored results on Noometry and Mistral Small 3.1 has 28.

Related comparisons

Go deeper