Model comparison

Mistral Small vs o1

o1 is the stronger model overall, scoring 40.9 to 33.4 on the Noometry Index. Mistral Small costs 100× less per token, which makes it the better buy when o1's lead doesn't matter for your workload.

Last verified . 30 shared benchmarks.

Mistral Small Mistral AI

33.4

Rank #243 Confirmed

o1 OpenAI

40.9

Rank #143 Confirmed

Summary

  • They share 30 benchmarks with published results for both. Mistral Small scores higher in 1 category and o1 in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where o1 leads 36.1 to 16.4.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 5.8% for Mistral Small and 73.3% for o1.
  • Mistral Small is cheaper at $0.15 / $0.60 per million input/output tokens, against $15 / $60 for o1.
  • Mistral Small accepts more context: 262K tokens versus 200K.
  • Mistral Small has downloadable open weights; the other is API-only.

Side by side

Mistral Small and o1 specifications
Mistral Smallo1
ProviderMistral AIOpenAI
Noometry Index33.440.9
Released2024-02-262024-09-12
WeightsOpenProprietary
Context window262K200K
Max output256K100K
Input $ / M tokens$0.15$15
Output $ / M tokens$0.60$60
Results tracked3952

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1 leads

Mistral Small: 34.0 (#247), o1: 46.1 (#70)

Coding benchmarks
BenchmarkMistral Smallo1
LiveBench Coding36.2%69.7%
LMArena Coding13621367
Aider Polyglot—61.7%
SciCode26.5%—
WeirdML—47.6%
BigCodeBench Instruct36.1%—
BigCodeBench Complete46.6%—
CadEval—56%
ALE-Bench497.62—
HumanEval+—89%
MBPP+—80.2%

Agentic & Tool Use Mistral Small leads

Mistral Small: 28.1 (#93), o1: 24.6 (#117)

Agentic & Tool Use benchmarks
BenchmarkMistral Smallo1
Berkeley Function Calling Leaderboard37.1%—
Cybench—10%
METR Time Horizons—51.1%

Reasoning o1 leads

Mistral Small: 19.8 (#250), o1: 27.9 (#111)

Reasoning benchmarks
BenchmarkMistral Smallo1
LiveBench Reasoning44.8%91.6%
LMArena Hard Prompts13351371
DTBench70.9%74.7%
LiveBench Data Analysis53.7%65.5%
LMCA20.6%22.3%
LiveBench44%75.7%
SimpleBench—41.7%
Kagi LLM Benchmark37.8%—
ARC-AGI-1—30.7%
CritPt0%—
Chess Puzzles—15%
EnigmaEval—5.7%
Epoch Capabilities Index—141.91

Math o1 leads

Mistral Small: 16.4 (#293), o1: 36.1 (#175)

Math benchmarks
BenchmarkMistral Smallo1
OTIS Mock AIME 2024-20255.8%73.3%
LiveBench Math39.9%80.3%
LMArena Math13411388
MATH Level 546.8%94.7%
FrontierMath (Tiers 1-3)—14.7%
FrontierMath (Feb 2025 set)—9.3%

Knowledge o1 leads

Mistral Small: 31.0 (#222), o1: 41.5 (#110)

Knowledge benchmarks
BenchmarkMistral Smallo1
GPQA Diamond47.5%76.8%
LMArena Expert12911361
Humanity's Last Exam—8%
SimpleQA Verified—41.1%
Confabulations—11.7%
Vectara Hallucination Rate5.1%—
MMLU68.7%—

Multimodal Too close to call

Mistral Small: 33.5 (#96), o1: 34.2 (#93)

Multimodal benchmarks
BenchmarkMistral Smallo1
LMArena Vision11421168
GeoBench—80%
VPCT—37%
SpatialViz-Bench—41.4%

Multilingual o1 leads

Mistral Small: 45.5 (#169), o1: 48.6 (#142)

Multilingual benchmarks
BenchmarkMistral Smallo1
LMArena Non-English13151358
LMArena Chinese13401394
LMArena French13371344
LMArena German13401337
LMArena Japanese12751346
LMArena Korean12591396
LMArena Russian13241356
LMArena Spanish13461345

Instruction Following o1 leads

Mistral Small: 66.4 (#209), o1: 74.8 (#86)

Instruction Following benchmarks
BenchmarkMistral Smallo1
LiveBench Instruction Following63.7%81.5%
LMArena Instruction Following13101367

Long Context o1 leads

Mistral Small: 40.4 (#156), o1: 50.3 (#9)

Long Context benchmarks
BenchmarkMistral Smallo1
LMArena Longer Query13271378
Fiction.LiveBench—83.3%

Writing & Preference o1 leads

Mistral Small: 52.5 (#171), o1: 55.6 (#144)

Writing & Preference benchmarks
BenchmarkMistral Smallo1
LMArena Text13381366
LMArena Creative Writing13051348
LMArena Multi-Turn13441369
LiveBench Language30.5%65.4%
Short-Story Creative Writing—70.2%

Frequently asked questions

Is Mistral Small better than o1?

o1 is the stronger model overall, scoring 40.9 to 33.4 on the Noometry Index. Mistral Small costs 100× less per token, which makes it the better buy when o1's lead doesn't matter for your workload.

Which is cheaper, Mistral Small or o1?

Mistral Small is cheaper. It lists at $0.15 per million input tokens and $0.60 per million output tokens; o1 lists at $15 and $60.

Is Mistral Small or o1 better for coding?

o1 scores higher on coding benchmarks: 46.1 versus 34.0 in the Noometry coding category.

Which has the bigger context window?

Mistral Small does, with 262K tokens against 200K.

How many benchmarks do Mistral Small and o1 share?

30 benchmarks have published results for both models. Mistral Small has 39 scored results on Noometry and o1 has 52.

Related comparisons

Go deeper