Model comparison

Mistral Small 3.1 vs o1

o1 is the stronger model overall, scoring 40.9 to 31.7 on the Noometry Index. Mistral Small 3.1 costs 65× less per token, which makes it the better buy when o1's lead doesn't matter for your workload.

Last verified . 22 shared benchmarks.

Mistral Small 3.1 Mistral AI

31.7

Rank #269 Confirmed

o1 OpenAI

40.9

Rank #143 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Mistral Small 3.1 scores higher in 0 categories and o1 in 9 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where o1 leads 36.1 to 14.7.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 3.9% for Mistral Small 3.1 and 73.3% for o1.
  • Mistral Small 3.1 is cheaper at $0.35 / $0.56 per million input/output tokens, against $15 / $60 for o1.
  • o1 accepts more context: 200K tokens versus 128K.
  • Mistral Small 3.1 has downloadable open weights; the other is API-only.

Side by side

Mistral Small 3.1 and o1 specifications
Mistral Small 3.1o1
ProviderMistral AIOpenAI
Noometry Index31.740.9
Released2025-03-172024-09-12
WeightsOpenProprietary
Context window128K200K
Max output102K100K
Input $ / M tokens$0.35$15
Output $ / M tokens$0.56$60
Results tracked2852

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1 leads

Mistral Small 3.1: 38.3 (#179), o1: 46.1 (#70)

Coding benchmarks
BenchmarkMistral Small 3.1o1
LMArena Coding13091367
Aider Polyglot—61.7%
WeirdML—47.6%
LiveBench Coding—69.7%
CadEval—56%
HumanEval+—89%
MBPP+—80.2%

Agentic & Tool Use Not comparable

Mistral Small 3.1: —, o1: 24.6 (#117)

Agentic & Tool Use benchmarks
BenchmarkMistral Small 3.1o1
Cybench—10%
METR Time Horizons—51.1%

Reasoning o1 leads

Mistral Small 3.1: 19.7 (#254), o1: 27.9 (#111)

Reasoning benchmarks
BenchmarkMistral Small 3.1o1
Chess Puzzles1%15%
LMArena Hard Prompts12781371
Epoch Capabilities Index127.48141.91
SimpleBench—41.7%
ARC-AGI-1—30.7%
EnigmaEval—5.7%
LiveBench Reasoning—91.6%
DTBench—74.7%
LiveBench Data Analysis—65.5%
LMCA—22.3%
LiveBench—75.7%

Math o1 leads

Mistral Small 3.1: 14.7 (#301), o1: 36.1 (#175)

Math benchmarks
BenchmarkMistral Small 3.1o1
OTIS Mock AIME 2024-20253.9%73.3%
LMArena Math12621388
FrontierMath (Tiers 1-3)—14.7%
Omni-MATH24.8%—
LiveBench Math—80.3%
MATH Level 5—94.7%
FrontierMath (Feb 2025 set)—9.3%

Knowledge o1 leads

Mistral Small 3.1: 22.6 (#271), o1: 41.5 (#110)

Knowledge benchmarks
BenchmarkMistral Small 3.1o1
GPQA Diamond41.9%76.8%
LMArena Expert12571361
Humanity's Last Exam—8%
SimpleQA Verified—41.1%
MMLU-Pro61%—
Confabulations—11.7%
GPQA (HELM)39.2%—

Multimodal Too close to call

Mistral Small 3.1: 33.2 (#99), o1: 34.2 (#93)

Multimodal benchmarks
BenchmarkMistral Small 3.1o1
LMArena Vision11361168
GeoBench—80%
VPCT—37%
SpatialViz-Bench—41.4%

Multilingual o1 leads

Mistral Small 3.1: 41.2 (#209), o1: 48.6 (#142)

Multilingual benchmarks
BenchmarkMistral Small 3.1o1
LMArena Non-English12551358
LMArena Chinese12531394
LMArena French12731344
LMArena German12661337
LMArena Japanese12081346
LMArena Korean12061396
LMArena Russian12631356
LMArena Spanish12831345

Instruction Following o1 leads

Mistral Small 3.1: 63.6 (#230), o1: 74.8 (#86)

Instruction Following benchmarks
BenchmarkMistral Small 3.1o1
LMArena Instruction Following12641367
LiveBench Instruction Following—81.5%
IFEval75%—

Long Context o1 leads

Mistral Small 3.1: 39.5 (#178), o1: 50.3 (#9)

Long Context benchmarks
BenchmarkMistral Small 3.1o1
LMArena Longer Query12991378
Fiction.LiveBench—83.3%

Writing & Preference o1 leads

Mistral Small 3.1: 37.0 (#259), o1: 55.6 (#144)

Writing & Preference benchmarks
BenchmarkMistral Small 3.1o1
LMArena Text12771366
LMArena Creative Writing12531348
LMArena Multi-Turn12701369
Short-Story Creative Writing—70.2%
EQ-Bench Creative Writing761—
WildBench78.8%—
LiveBench Language—65.4%

Frequently asked questions

Is Mistral Small 3.1 better than o1?

o1 is the stronger model overall, scoring 40.9 to 31.7 on the Noometry Index. Mistral Small 3.1 costs 65× less per token, which makes it the better buy when o1's lead doesn't matter for your workload.

Which is cheaper, Mistral Small 3.1 or o1?

Mistral Small 3.1 is cheaper. It lists at $0.35 per million input tokens and $0.56 per million output tokens; o1 lists at $15 and $60.

Is Mistral Small 3.1 or o1 better for coding?

o1 scores higher on coding benchmarks: 46.1 versus 38.3 in the Noometry coding category.

Which has the bigger context window?

o1 does, with 200K tokens against 128K.

How many benchmarks do Mistral Small 3.1 and o1 share?

22 benchmarks have published results for both models. Mistral Small 3.1 has 28 scored results on Noometry and o1 has 52.

Related comparisons

Go deeper