Model comparison

Mistral Medium 3.1 vs o1

o1 is the stronger model overall, scoring 40.9 to 31.9 on the Noometry Index. Mistral Medium 3.1 costs 33× less per token, which makes it the better buy when o1's lead doesn't matter for your workload.

Last verified . 0 shared benchmarks.

Mistral Medium 3.1 Mistral AI

31.9

Rank #266 Reported

o1 OpenAI

40.9

Rank #143 Confirmed

Summary

  • The widest gap is in reasoning, where o1 leads 27.9 to 10.6.
  • Mistral Medium 3.1 is cheaper at $0.40 / $2 per million input/output tokens, against $15 / $60 for o1.
  • o1 accepts more context: 200K tokens versus 131K.

Side by side

Mistral Medium 3.1 and o1 specifications
Mistral Medium 3.1o1
ProviderMistral AIOpenAI
Noometry Index31.940.9
Released—2024-09-12
WeightsProprietaryProprietary
Context window131K200K
Max output105K100K
Input $ / M tokens$0.40$15
Output $ / M tokens$2$60
Results tracked352

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Mistral Medium 3.1: —, o1: 46.1 (#70)

Coding benchmarks
BenchmarkMistral Medium 3.1o1
Aider Polyglot—61.7%
WeirdML—47.6%
LiveBench Coding—69.7%
LMArena Coding—1367
CadEval—56%
HumanEval+—89%
MBPP+—80.2%

Agentic & Tool Use Not comparable

Mistral Medium 3.1: —, o1: 24.6 (#117)

Agentic & Tool Use benchmarks
BenchmarkMistral Medium 3.1o1
Cybench—10%
METR Time Horizons—51.1%

Reasoning o1 leads

Mistral Medium 3.1: 10.6 (#341), o1: 27.9 (#111)

Reasoning benchmarks
BenchmarkMistral Medium 3.1o1
SimpleBench—41.7%
NYT Connections (extended)6.5%—
ARC-AGI-1—30.7%
Chess Puzzles—15%
EnigmaEval—5.7%
Thematic Generalization20.3%—
LiveBench Reasoning—91.6%
LMArena Hard Prompts—1371
DTBench—74.7%
LiveBench Data Analysis—65.5%
LMCA—22.3%
Epoch Capabilities Index—141.91
LiveBench—75.7%

Math Not comparable

Mistral Medium 3.1: —, o1: 36.1 (#175)

Math benchmarks
BenchmarkMistral Medium 3.1o1
FrontierMath (Tiers 1-3)—14.7%
OTIS Mock AIME 2024-2025—73.3%
LiveBench Math—80.3%
LMArena Math—1388
MATH Level 5—94.7%
FrontierMath (Feb 2025 set)—9.3%

Knowledge Not comparable

Mistral Medium 3.1: —, o1: 41.5 (#110)

Knowledge benchmarks
BenchmarkMistral Medium 3.1o1
GPQA Diamond—76.8%
Humanity's Last Exam—8%
SimpleQA Verified—41.1%
Confabulations—11.7%
LMArena Expert—1361

Multimodal Not comparable

Mistral Medium 3.1: —, o1: 34.2 (#93)

Multimodal benchmarks
BenchmarkMistral Medium 3.1o1
LMArena Vision—1168
GeoBench—80%
VPCT—37%
SpatialViz-Bench—41.4%

Multilingual Not comparable

Mistral Medium 3.1: —, o1: 48.6 (#142)

Multilingual benchmarks
BenchmarkMistral Medium 3.1o1
LMArena Non-English—1358
LMArena Chinese—1394
LMArena French—1344
LMArena German—1337
LMArena Japanese—1346
LMArena Korean—1396
LMArena Russian—1356
LMArena Spanish—1345

Instruction Following Not comparable

Mistral Medium 3.1: —, o1: 74.8 (#86)

Instruction Following benchmarks
BenchmarkMistral Medium 3.1o1
LiveBench Instruction Following—81.5%
LMArena Instruction Following—1367

Long Context Not comparable

Mistral Medium 3.1: —, o1: 50.3 (#9)

Long Context benchmarks
BenchmarkMistral Medium 3.1o1
Fiction.LiveBench—83.3%
LMArena Longer Query—1378

Writing & Preference Too close to call

Mistral Medium 3.1: 55.5 (#145), o1: 55.6 (#144)

Writing & Preference benchmarks
BenchmarkMistral Medium 3.1o1
LMArena Text—1366
LMArena Creative Writing—1348
Short-Story Creative Writing—70.2%
EQ-Bench Creative Writing1476—
LMArena Multi-Turn—1369
LiveBench Language—65.4%

Frequently asked questions

Is Mistral Medium 3.1 better than o1?

o1 is the stronger model overall, scoring 40.9 to 31.9 on the Noometry Index. Mistral Medium 3.1 costs 33× less per token, which makes it the better buy when o1's lead doesn't matter for your workload.

Which is cheaper, Mistral Medium 3.1 or o1?

Mistral Medium 3.1 is cheaper. It lists at $0.40 per million input tokens and $2 per million output tokens; o1 lists at $15 and $60.

Which has the bigger context window?

o1 does, with 200K tokens against 131K.

How many benchmarks do Mistral Medium 3.1 and o1 share?

0 benchmarks have published results for both models. Mistral Medium 3.1 has 3 scored results on Noometry and o1 has 52.

Related comparisons

Go deeper