Model comparison

Mistral Medium vs o1

o1 is the stronger model overall, scoring 40.9 to 36.3 on the Noometry Index. Mistral Medium costs 8.8× less per token, which makes it the better buy when o1's lead doesn't matter for your workload.

Last verified . 27 shared benchmarks.

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

o1 OpenAI

40.9

Rank #143 Confirmed

Summary

  • They share 27 benchmarks with published results for both. Mistral Medium scores higher in 4 categories and o1 in 6 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where o1 leads 41.5 to 25.0.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 32.2% for Mistral Medium and 73.3% for o1.
  • Mistral Medium is cheaper at $1.50 / $7.50 per million input/output tokens, against $15 / $60 for o1.
  • Mistral Medium accepts more context: 262K tokens versus 200K.
  • Mistral Medium has downloadable open weights; the other is API-only.

Side by side

Mistral Medium and o1 specifications
Mistral Mediumo1
ProviderMistral AIOpenAI
Noometry Index36.340.9
Released2023-12-112024-09-12
WeightsOpenProprietary
Context window262K200K
Max output262K100K
Input $ / M tokens$1.50$15
Output $ / M tokens$7.50$60
Results tracked3652

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1 leads

Mistral Medium: 34.2 (#243), o1: 46.1 (#70)

Coding benchmarks
BenchmarkMistral Mediumo1
WeirdML43.7%47.6%
LMArena Coding14341367
FrontierCode8%—
Aider Polyglot—61.7%
SciCode40.2%—
LiveBench Coding—69.7%
CadEval—56%
ALE-Bench763.98—
HumanEval+—89%
MBPP+—80.2%

Agentic & Tool Use Mistral Medium leads

Mistral Medium: 28.3 (#90), o1: 24.6 (#117)

Agentic & Tool Use benchmarks
BenchmarkMistral Mediumo1
Berkeley Function Calling Leaderboard37.7%—
Cybench—10%
METR Time Horizons—51.1%

Reasoning o1 leads

Mistral Medium: 24.0 (#167), o1: 27.9 (#111)

Reasoning benchmarks
BenchmarkMistral Mediumo1
LMArena Hard Prompts14261371
DTBench75.5%74.7%
LMCA26.1%22.3%
SimpleBench—41.7%
Kagi LLM Benchmark50%—
ARC-AGI-1—30.7%
CritPt0%—
Chess Puzzles—15%
EnigmaEval—5.7%
LiveBench Reasoning—91.6%
LiveBench Data Analysis—65.5%
Surface Evolver Bench26.9%—
Epoch Capabilities Index—141.91
LiveBench—75.7%

Math o1 leads

Mistral Medium: 28.1 (#245), o1: 36.1 (#175)

Math benchmarks
BenchmarkMistral Mediumo1
OTIS Mock AIME 2024-202532.2%73.3%
LMArena Math14081388
MATH Level 581.6%94.7%
FrontierMath (Feb 2025 set)0.3%9.3%
FrontierMath (Tiers 1-3)—14.7%
ProofBench9%—
LiveBench Math—80.3%

Knowledge o1 leads

Mistral Medium: 25.0 (#265), o1: 41.5 (#110)

Knowledge benchmarks
BenchmarkMistral Mediumo1
GPQA Diamond59.5%76.8%
Humanity's Last Exam4.5%8%
LMArena Expert14081361
SimpleQA Verified—41.1%
Confabulations—11.7%
Vectara Hallucination Rate22.7%—

Multimodal Mistral Medium leads

Mistral Medium: 35.3 (#88), o1: 34.2 (#93)

Multimodal benchmarks
BenchmarkMistral Mediumo1
LMArena Vision11721168
GeoBench—80%
VPCT—37%
SpatialViz-Bench—41.4%

Multilingual Mistral Medium leads

Mistral Medium: 52.1 (#91), o1: 48.6 (#142)

Multilingual benchmarks
BenchmarkMistral Mediumo1
LMArena Non-English14081358
LMArena Chinese14471394
LMArena French14591344
LMArena German14321337
LMArena Japanese13781346
LMArena Korean13801396
LMArena Russian14111356
LMArena Spanish14331345

Instruction Following o1 leads

Mistral Medium: 73.7 (#116), o1: 74.8 (#86)

Instruction Following benchmarks
BenchmarkMistral Mediumo1
LMArena Instruction Following13981367
LiveBench Instruction Following—81.5%

Long Context o1 leads

Mistral Medium: 42.9 (#114), o1: 50.3 (#9)

Long Context benchmarks
BenchmarkMistral Mediumo1
LMArena Longer Query14061378
Fiction.LiveBench—83.3%

Writing & Preference Mistral Medium leads

Mistral Medium: 60.0 (#103), o1: 55.6 (#144)

Writing & Preference benchmarks
BenchmarkMistral Mediumo1
LMArena Text14241366
LMArena Creative Writing13911348
Short-Story Creative Writing77.3%70.2%
LMArena Multi-Turn14181369
LiveBench Language—65.4%

Frequently asked questions

Is Mistral Medium better than o1?

o1 is the stronger model overall, scoring 40.9 to 36.3 on the Noometry Index. Mistral Medium costs 8.8× less per token, which makes it the better buy when o1's lead doesn't matter for your workload.

Which is cheaper, Mistral Medium or o1?

Mistral Medium is cheaper. It lists at $1.50 per million input tokens and $7.50 per million output tokens; o1 lists at $15 and $60.

Is Mistral Medium or o1 better for coding?

o1 scores higher on coding benchmarks: 46.1 versus 34.2 in the Noometry coding category.

Which has the bigger context window?

Mistral Medium does, with 262K tokens against 200K.

How many benchmarks do Mistral Medium and o1 share?

27 benchmarks have published results for both models. Mistral Medium has 36 scored results on Noometry and o1 has 52.

Related comparisons

Go deeper