Model comparison

Mistral Small 3.1 vs Olmo 3.1 32b Instruct

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 31.7 on the Noometry Index.

Last verified . 16 shared benchmarks.

Summary

  • They share 16 benchmarks with published results for both. Mistral Small 3.1 scores higher in 0 categories and Olmo 3.1 32b Instruct in 8 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Olmo 3.1 32b Instruct leads 36.3 to 14.7.

Side by side

Mistral Small 3.1 and Olmo 3.1 32b Instruct specifications
Mistral Small 3.1Olmo 3.1 32b Instruct
ProviderMistral AIAllen Institute for AI (Ai2)
Noometry Index31.739.4
Released2025-03-17—
WeightsOpenOpen
Context window128K—
Max output102K—
Input $ / M tokens$0.35—
Output $ / M tokens$0.56—
Results tracked2816

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Instruct leads

Mistral Small 3.1: 38.3 (#179), Olmo 3.1 32b Instruct: 39.5 (#157)

Coding benchmarks
BenchmarkMistral Small 3.1Olmo 3.1 32b Instruct
LMArena Coding13091347

Reasoning Olmo 3.1 32b Instruct leads

Mistral Small 3.1: 19.7 (#254), Olmo 3.1 32b Instruct: 26.4 (#132)

Reasoning benchmarks
BenchmarkMistral Small 3.1Olmo 3.1 32b Instruct
LMArena Hard Prompts12781322
Chess Puzzles1%—
Epoch Capabilities Index127.48—

Math Olmo 3.1 32b Instruct leads

Mistral Small 3.1: 14.7 (#301), Olmo 3.1 32b Instruct: 36.3 (#167)

Math benchmarks
BenchmarkMistral Small 3.1Olmo 3.1 32b Instruct
LMArena Math12621305
OTIS Mock AIME 2024-20253.9%—
Omni-MATH24.8%—

Knowledge Olmo 3.1 32b Instruct leads

Mistral Small 3.1: 22.6 (#271), Olmo 3.1 32b Instruct: 36.1 (#175)

Knowledge benchmarks
BenchmarkMistral Small 3.1Olmo 3.1 32b Instruct
LMArena Expert12571308
GPQA Diamond41.9%—
MMLU-Pro61%—
GPQA (HELM)39.2%—

Multimodal Not comparable

Mistral Small 3.1: 33.2 (#99), Olmo 3.1 32b Instruct: —

Multimodal benchmarks
BenchmarkMistral Small 3.1Olmo 3.1 32b Instruct
LMArena Vision1136—

Multilingual Olmo 3.1 32b Instruct leads

Mistral Small 3.1: 41.2 (#209), Olmo 3.1 32b Instruct: 42.6 (#191)

Multilingual benchmarks
BenchmarkMistral Small 3.1Olmo 3.1 32b Instruct
LMArena Non-English12551275
LMArena Chinese12531304
LMArena French12731328
LMArena German12661282
LMArena Korean12061206
LMArena Russian12631268
LMArena Spanish12831336
LMArena Japanese1208—

Instruction Following Olmo 3.1 32b Instruct leads

Mistral Small 3.1: 63.6 (#230), Olmo 3.1 32b Instruct: 68.6 (#187)

Instruction Following benchmarks
BenchmarkMistral Small 3.1Olmo 3.1 32b Instruct
LMArena Instruction Following12641299
IFEval75%—

Long Context Too close to call

Mistral Small 3.1: 39.5 (#178), Olmo 3.1 32b Instruct: 39.9 (#166)

Long Context benchmarks
BenchmarkMistral Small 3.1Olmo 3.1 32b Instruct
LMArena Longer Query12991312

Writing & Preference Olmo 3.1 32b Instruct leads

Mistral Small 3.1: 37.0 (#259), Olmo 3.1 32b Instruct: 50.2 (#185)

Writing & Preference benchmarks
BenchmarkMistral Small 3.1Olmo 3.1 32b Instruct
LMArena Text12771311
LMArena Creative Writing12531264
LMArena Multi-Turn12701309
EQ-Bench Creative Writing761—
WildBench78.8%—

Frequently asked questions

Is Mistral Small 3.1 better than Olmo 3.1 32b Instruct?

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 31.7 on the Noometry Index.

Is Mistral Small 3.1 or Olmo 3.1 32b Instruct better for coding?

Olmo 3.1 32b Instruct scores higher on coding benchmarks: 39.5 versus 38.3 in the Noometry coding category.

How many benchmarks do Mistral Small 3.1 and Olmo 3.1 32b Instruct share?

16 benchmarks have published results for both models. Mistral Small 3.1 has 28 scored results on Noometry and Olmo 3.1 32b Instruct has 16.

Related comparisons

Go deeper