Model comparison

Mistral Medium vs o3

o3 is the stronger model overall, scoring 47.5 to 36.3 on the Noometry Index.

Last verified . 31 shared benchmarks.

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 31 benchmarks with published results for both. Mistral Medium scores higher in 2 categories and o3 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where o3 leads 54.6 to 25.0.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 32.2% for Mistral Medium and 84.4% for o3.
  • Mistral Medium is cheaper at $1.50 / $7.50 per million input/output tokens, against $2 / $8 for o3.
  • Mistral Medium accepts more context: 262K tokens versus 200K.
  • Mistral Medium has downloadable open weights; the other is API-only.

Side by side

Mistral Medium and o3 specifications
Mistral Mediumo3
ProviderMistral AIOpenAI
Noometry Index36.347.5
Released2023-12-112025-04-16
WeightsOpenProprietary
Context window262K200K
Max output262K100K
Input $ / M tokens$1.50$2
Output $ / M tokens$7.50$8
Results tracked3663

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3 leads

Mistral Medium: 34.2 (#243), o3: 46.8 (#64)

Coding benchmarks
BenchmarkMistral Mediumo3
WeirdML43.7%52.4%
LMArena Coding14341408
ALE-Bench763.98933.55
SWE-bench Verified—62.3%
FrontierCode8%—
SWE-bench Verified (bash only)—58.4%
Aider Polyglot—81.3%
SciCode40.2%—
GSO—8.8%
CadEval—74%

Agentic & Tool Use o3 leads

Mistral Medium: 28.3 (#90), o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkMistral Mediumo3
Berkeley Function Calling Leaderboard37.7%63%
GDPval—30.8%
DeepResearch Bench—45.2%
OSWorld—23%
LMArena Search—1144
METR Time Horizons—65.4%

Reasoning o3 leads

Mistral Medium: 24.0 (#167), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkMistral Mediumo3
Kagi LLM Benchmark50%67.6%
CritPt0%1.4%
LMArena Hard Prompts14261402
DTBench75.5%84.8%
LMCA26.1%39.7%
ARC-AGI-2—6.5%
SimpleBench—53.1%
ARC-AGI-1—60.8%
Chess Puzzles—38%
EnigmaEval—13.1%
Mystery Game Puzzles—29%
Surface Evolver Bench26.9%—
Epoch Capabilities Index—146.86
ForecastBench—62.5

Math o3 leads

Mistral Medium: 28.1 (#245), o3: 50.2 (#58)

Math benchmarks
BenchmarkMistral Mediumo3
OTIS Mock AIME 2024-202532.2%84.4%
LMArena Math14081426
MATH Level 581.6%97.8%
FrontierMath (Feb 2025 set)0.3%18.7%
FrontierMath (Tiers 1-3)—33.3%
ProofBench9%—
Omni-MATH—71.4%
FrontierMath Tier 4 (v1)—2.1%

Knowledge o3 leads

Mistral Medium: 25.0 (#265), o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkMistral Mediumo3
GPQA Diamond59.5%81.8%
Humanity's Last Exam4.5%20.3%
LMArena Expert14081402
SimpleQA Verified—49.4%
MMLU-Pro—85.9%
Confabulations—14.4%
Vectara Hallucination Rate22.7%—
GPQA (HELM)—75.3%

Multimodal o3 leads

Mistral Medium: 35.3 (#88), o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkMistral Mediumo3
LMArena Vision11721214
GeoBench—74%
VPCT—52%

Multilingual Too close to call

Mistral Medium: 52.1 (#91), o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkMistral Mediumo3
LMArena Non-English14081401
LMArena Chinese14471437
LMArena French14591430
LMArena German14321420
LMArena Japanese13781403
LMArena Korean13801370
LMArena Russian14111406
LMArena Spanish14331395

Instruction Following Too close to call

Mistral Medium: 73.7 (#116), o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkMistral Mediumo3
LMArena Instruction Following13981368
IFEval—86.9%

Long Context o3 leads

Mistral Medium: 42.9 (#114), o3: 53.3 (#6)

Long Context benchmarks
BenchmarkMistral Mediumo3
LMArena Longer Query14061372
Fiction.LiveBench—88.9%
CL-bench—17.8%

Writing & Preference o3 leads

Mistral Medium: 60.0 (#103), o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkMistral Mediumo3
LMArena Text14241410
LMArena Creative Writing13911359
Short-Story Creative Writing77.3%83.9%
LMArena Multi-Turn14181405
EQ-Bench Creative Writing—1676
WildBench—86.1%

Frequently asked questions

Is Mistral Medium better than o3?

o3 is the stronger model overall, scoring 47.5 to 36.3 on the Noometry Index.

Which is cheaper, Mistral Medium or o3?

Mistral Medium is cheaper. It lists at $1.50 per million input tokens and $7.50 per million output tokens; o3 lists at $2 and $8.

Is Mistral Medium or o3 better for coding?

o3 scores higher on coding benchmarks: 46.8 versus 34.2 in the Noometry coding category.

Which has the bigger context window?

Mistral Medium does, with 262K tokens against 200K.

How many benchmarks do Mistral Medium and o3 share?

31 benchmarks have published results for both models. Mistral Medium has 36 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper