Model comparison

Mistral Small 3.1 vs Olmo 3.1 32b Think

Olmo 3.1 32b Think is the stronger model overall, scoring 37.9 to 31.7 on the Noometry Index.

Last verified . 15 shared benchmarks.

Mistral Small 3.1 Mistral AI

31.7

Rank #269 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Mistral Small 3.1 scores higher in 3 categories and Olmo 3.1 32b Think in 5 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where Olmo 3.1 32b Think leads 36.3 to 14.7.

Side by side

Mistral Small 3.1 and Olmo 3.1 32b Think specifications
Mistral Small 3.1Olmo 3.1 32b Think
ProviderMistral AIAllen Institute for AI (Ai2)
Noometry Index31.737.9
Released2025-03-17—
WeightsOpenOpen
Context window128K—
Max output102K—
Input $ / M tokens$0.35—
Output $ / M tokens$0.56—
Results tracked2815

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Mistral Small 3.1: 38.3 (#179), Olmo 3.1 32b Think: 37.7 (#189)

Coding benchmarks
BenchmarkMistral Small 3.1Olmo 3.1 32b Think
LMArena Coding13091291

Reasoning Olmo 3.1 32b Think leads

Mistral Small 3.1: 19.7 (#254), Olmo 3.1 32b Think: 25.2 (#150)

Reasoning benchmarks
BenchmarkMistral Small 3.1Olmo 3.1 32b Think
LMArena Hard Prompts12781272
Chess Puzzles1%—
Epoch Capabilities Index127.48—

Math Olmo 3.1 32b Think leads

Mistral Small 3.1: 14.7 (#301), Olmo 3.1 32b Think: 36.3 (#168)

Math benchmarks
BenchmarkMistral Small 3.1Olmo 3.1 32b Think
LMArena Math12621305
OTIS Mock AIME 2024-20253.9%—
Omni-MATH24.8%—

Knowledge Olmo 3.1 32b Think leads

Mistral Small 3.1: 22.6 (#271), Olmo 3.1 32b Think: 35.7 (#181)

Knowledge benchmarks
BenchmarkMistral Small 3.1Olmo 3.1 32b Think
LMArena Expert12571295
GPQA Diamond41.9%—
MMLU-Pro61%—
GPQA (HELM)39.2%—

Multimodal Not comparable

Mistral Small 3.1: 33.2 (#99), Olmo 3.1 32b Think: —

Multimodal benchmarks
BenchmarkMistral Small 3.1Olmo 3.1 32b Think
LMArena Vision1136—

Multilingual Mistral Small 3.1 leads

Mistral Small 3.1: 41.2 (#209), Olmo 3.1 32b Think: 38.1 (#231)

Multilingual benchmarks
BenchmarkMistral Small 3.1Olmo 3.1 32b Think
LMArena Non-English12551209
LMArena Chinese12531242
LMArena French12731260
LMArena German12661262
LMArena Russian12631193
LMArena Spanish12831289
LMArena Japanese1208—
LMArena Korean1206—

Instruction Following Olmo 3.1 32b Think leads

Mistral Small 3.1: 63.6 (#230), Olmo 3.1 32b Think: 65.6 (#218)

Instruction Following benchmarks
BenchmarkMistral Small 3.1Olmo 3.1 32b Think
LMArena Instruction Following12641247
IFEval75%—

Long Context Too close to call

Mistral Small 3.1: 39.5 (#178), Olmo 3.1 32b Think: 38.6 (#195)

Long Context benchmarks
BenchmarkMistral Small 3.1Olmo 3.1 32b Think
LMArena Longer Query12991272

Writing & Preference Olmo 3.1 32b Think leads

Mistral Small 3.1: 37.0 (#259), Olmo 3.1 32b Think: 46.2 (#220)

Writing & Preference benchmarks
BenchmarkMistral Small 3.1Olmo 3.1 32b Think
LMArena Text12771272
LMArena Creative Writing12531226
LMArena Multi-Turn12701252
EQ-Bench Creative Writing761—
WildBench78.8%—

Frequently asked questions

Is Mistral Small 3.1 better than Olmo 3.1 32b Think?

Olmo 3.1 32b Think is the stronger model overall, scoring 37.9 to 31.7 on the Noometry Index.

Is Mistral Small 3.1 or Olmo 3.1 32b Think better for coding?

They score almost the same on coding (38.3 vs 37.7); test both on your own repository before choosing.

How many benchmarks do Mistral Small 3.1 and Olmo 3.1 32b Think share?

15 benchmarks have published results for both models. Mistral Small 3.1 has 28 scored results on Noometry and Olmo 3.1 32b Think has 15.

Related comparisons

Go deeper