Model comparison

Mistral vs Mistral Small 3

Mistral Small 3 is the stronger model overall, scoring 31.2 to 29.9 on the Noometry Index.

Last verified . 16 shared benchmarks.

Mistral Mistral AI

29.9

Rank #303 Confirmed

Mistral Small 3 Mistral AI

31.2

Rank #278 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Mistral scores higher in 3 categories and Mistral Small 3 in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Mistral Small 3 leads 63.7 to 52.6.
  • Mistral Small 3 has downloadable open weights; the other is API-only.

Side by side

Mistral and Mistral Small 3 specifications
MistralMistral Small 3
ProviderMistral AIMistral AI
Noometry Index29.931.2
Released—2025-01-30
WeightsProprietaryOpen
Context window—33K
Max output—16K
Input $ / M tokens—$0.05
Output $ / M tokens—$0.08
Results tracked2224

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Small 3 leads

Mistral: 33.8 (#250), Mistral Small 3: 36.5 (#207)

Coding benchmarks
BenchmarkMistralMistral Small 3
LMArena Coding11621246
BigCodeBench Instruct—45.3%
BigCodeBench Complete—50.4%

Reasoning Mistral leads

Mistral: 22.2 (#200), Mistral Small 3: 18.9 (#273)

Reasoning benchmarks
BenchmarkMistralMistral Small 3
LMArena Hard Prompts11491233
Chess Puzzles—0%
Epoch Capabilities Index—127.07

Math Mistral leads

Mistral: 22.3 (#278), Mistral Small 3: 16.3 (#295)

Math benchmarks
BenchmarkMistralMistral Small 3
LMArena Math11801240
OTIS Mock AIME 2024-2025—6.7%
Omni-MATH7.2%—

Knowledge Mistral Small 3 leads

Mistral: 16.6 (#288), Mistral Small 3: 25.1 (#263)

Knowledge benchmarks
BenchmarkMistralMistral Small 3
LMArena Expert11251202
GPQA Diamond—47.3%
MMLU-Pro27.7%—
Confabulations—25.2%
GPQA (HELM)30.3%—

Multilingual Mistral Small 3 leads

Mistral: 32.8 (#254), Mistral Small 3: 37.3 (#236)

Multilingual benchmarks
BenchmarkMistralMistral Small 3
LMArena Non-English11291198
LMArena Chinese11091204
LMArena French11801203
LMArena German11551211
LMArena Japanese10131111
LMArena Korean10321188
LMArena Russian11681216
LMArena Spanish1143—

Instruction Following Mistral Small 3 leads

Mistral: 52.6 (#288), Mistral Small 3: 63.7 (#229)

Instruction Following benchmarks
BenchmarkMistralMistral Small 3
LMArena Instruction Following11521214
IFEval56.8%—

Long Context Mistral Small 3 leads

Mistral: 35.0 (#245), Mistral Small 3: 37.8 (#211)

Long Context benchmarks
BenchmarkMistralMistral Small 3
LMArena Longer Query11531246

Writing & Preference Mistral leads

Mistral: 37.0 (#260), Mistral Small 3: 32.2 (#280)

Writing & Preference benchmarks
BenchmarkMistralMistral Small 3
LMArena Text11651234
LMArena Creative Writing11581195
LMArena Multi-Turn11471217
EQ-Bench Creative Writing—707
WildBench66%—

Frequently asked questions

Is Mistral better than Mistral Small 3?

Mistral Small 3 is the stronger model overall, scoring 31.2 to 29.9 on the Noometry Index.

Is Mistral or Mistral Small 3 better for coding?

Mistral Small 3 scores higher on coding benchmarks: 36.5 versus 33.8 in the Noometry coding category.

How many benchmarks do Mistral and Mistral Small 3 share?

16 benchmarks have published results for both models. Mistral has 22 scored results on Noometry and Mistral Small 3 has 24.

Related comparisons

Go deeper