Model comparison

Mistral Medium 3.5 vs o1

Mistral Medium 3.5 and o1 score almost the same on the Noometry Index (40.2 vs 40.9), so choose on price, context window or the category you care about most.

Last verified . 18 shared benchmarks.

Mistral Medium 3.5 Mistral AI

40.2

Rank #152 Confirmed

o1 OpenAI

40.9

Rank #143 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Mistral Medium 3.5 scores higher in 4 categories and o1 in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where o1 leads 27.9 to 17.3.
  • Mistral Medium 3.5 is cheaper at $1.50 / $7.50 per million input/output tokens, against $15 / $60 for o1.
  • Mistral Medium 3.5 accepts more context: 262K tokens versus 200K.
  • Mistral Medium 3.5 has downloadable open weights; the other is API-only.

Side by side

Mistral Medium 3.5 and o1 specifications
Mistral Medium 3.5o1
ProviderMistral AIOpenAI
Noometry Index40.240.9
Released—2024-09-12
WeightsOpenProprietary
Context window262K200K
Max output210K100K
Input $ / M tokens$1.50$15
Output $ / M tokens$7.50$60
Results tracked2252

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1 leads

Mistral Medium 3.5: 36.0 (#213), o1: 46.1 (#70)

Coding benchmarks
BenchmarkMistral Medium 3.5o1
LMArena Coding14611367
Aider Polyglot—61.7%
LMArena WebDev1264—
WeirdML—47.6%
LiveBench Coding—69.7%
CadEval—56%
HumanEval+—89%
MBPP+—80.2%

Agentic & Tool Use Not comparable

Mistral Medium 3.5: —, o1: 24.6 (#117)

Agentic & Tool Use benchmarks
BenchmarkMistral Medium 3.5o1
Cybench—10%
METR Time Horizons—51.1%

Reasoning o1 leads

Mistral Medium 3.5: 17.3 (#295), o1: 27.9 (#111)

Reasoning benchmarks
BenchmarkMistral Medium 3.5o1
LMArena Hard Prompts14361371
Epoch Capabilities Index141.35141.91
SimpleBench—41.7%
Kagi LLM Benchmark41.4%—
NYT Connections (extended)12.9%—
ARC-AGI-1—30.7%
Chess Puzzles—15%
EnigmaEval—5.7%
LiveBench Reasoning—91.6%
DTBench—74.7%
LiveBench Data Analysis—65.5%
LMCA—22.3%
LiveBench—75.7%

Math Mistral Medium 3.5 leads

Mistral Medium 3.5: 39.1 (#113), o1: 36.1 (#175)

Math benchmarks
BenchmarkMistral Medium 3.5o1
LMArena Math14311388
FrontierMath (Tiers 1-3)—14.7%
OTIS Mock AIME 2024-2025—73.3%
LiveBench Math—80.3%
MATH Level 5—94.7%
FrontierMath (Feb 2025 set)—9.3%

Knowledge o1 leads

Mistral Medium 3.5: 40.0 (#126), o1: 41.5 (#110)

Knowledge benchmarks
BenchmarkMistral Medium 3.5o1
LMArena Expert14321361
GPQA Diamond—76.8%
Humanity's Last Exam—8%
SimpleQA Verified—41.1%
Confabulations—11.7%

Multimodal Mistral Medium 3.5 leads

Mistral Medium 3.5: 38.3 (#65), o1: 34.2 (#93)

Multimodal benchmarks
BenchmarkMistral Medium 3.5o1
LMArena Vision12231168
GeoBench—80%
VPCT—37%
SpatialViz-Bench—41.4%

Multilingual Mistral Medium 3.5 leads

Mistral Medium 3.5: 51.9 (#100), o1: 48.6 (#142)

Multilingual benchmarks
BenchmarkMistral Medium 3.5o1
LMArena Non-English14041358
LMArena Chinese14421394
LMArena French14481344
LMArena German14511337
LMArena Korean13851396
LMArena Russian13951356
LMArena Spanish14091345
LMArena Japanese—1346

Instruction Following Too close to call

Mistral Medium 3.5: 74.6 (#90), o1: 74.8 (#86)

Instruction Following benchmarks
BenchmarkMistral Medium 3.5o1
LMArena Instruction Following14151367
LiveBench Instruction Following—81.5%

Long Context o1 leads

Mistral Medium 3.5: 43.2 (#103), o1: 50.3 (#9)

Long Context benchmarks
BenchmarkMistral Medium 3.5o1
LMArena Longer Query14151378
Fiction.LiveBench—83.3%

Writing & Preference Mistral Medium 3.5 leads

Mistral Medium 3.5: 58.5 (#117), o1: 55.6 (#144)

Writing & Preference benchmarks
BenchmarkMistral Medium 3.5o1
LMArena Text14211366
LMArena Creative Writing13741348
LMArena Multi-Turn14231369
Short-Story Creative Writing—70.2%
EQ-Bench 4993—
LiveBench Language—65.4%

Frequently asked questions

Is Mistral Medium 3.5 better than o1?

Mistral Medium 3.5 and o1 score almost the same on the Noometry Index (40.2 vs 40.9), so choose on price, context window or the category you care about most.

Which is cheaper, Mistral Medium 3.5 or o1?

Mistral Medium 3.5 is cheaper. It lists at $1.50 per million input tokens and $7.50 per million output tokens; o1 lists at $15 and $60.

Is Mistral Medium 3.5 or o1 better for coding?

o1 scores higher on coding benchmarks: 46.1 versus 36.0 in the Noometry coding category.

Which has the bigger context window?

Mistral Medium 3.5 does, with 262K tokens against 200K.

How many benchmarks do Mistral Medium 3.5 and o1 share?

18 benchmarks have published results for both models. Mistral Medium 3.5 has 22 scored results on Noometry and o1 has 52.

Related comparisons

Go deeper