Model comparison

Mistral Large vs o3

o3 is the stronger model overall, scoring 47.5 to 31.9 on the Noometry Index.

Last verified . 37 shared benchmarks.

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 37 benchmarks with published results for both. Mistral Large scores higher in 0 categories and o3 in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where o3 leads 50.2 to 18.2.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 8.5% for Mistral Large and 84.4% for o3.
  • Mistral Large is cheaper at $2 / $6 per million input/output tokens, against $2 / $8 for o3.
  • o3 accepts more context: 200K tokens versus 131K.
  • Mistral Large has downloadable open weights; the other is API-only.

Side by side

Mistral Large and o3 specifications
Mistral Largeo3
ProviderMistral AIOpenAI
Noometry Index31.947.5
Released2024-02-262025-04-16
WeightsOpenProprietary
Context window131K200K
Max output16K100K
Input $ / M tokens$2$2
Output $ / M tokens$6$8
Results tracked5163

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3 leads

Mistral Large: 34.3 (#240), o3: 46.8 (#64)

Coding benchmarks
BenchmarkMistral Largeo3
LMArena Coding12771408
ALE-Bench264.7933.55
SWE-bench Verified—62.3%
SWE-bench Verified (bash only)—58.4%
Aider Polyglot—81.3%
SciCode36.2%—
GSO—8.8%
WeirdML—52.4%
BigCodeBench Instruct30%—
LiveBench Coding47.1%—
BigCodeBench Complete38.3%—
CadEval—74%
HumanEval+62.2%—
MBPP+59.5%—

Agentic & Tool Use o3 leads

Mistral Large: 28.6 (#89), o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkMistral Largeo3
Berkeley Function Calling Leaderboard38.4%63%
GDPval—30.8%
DeepResearch Bench—45.2%
OSWorld—23%
LMArena Search—1144
METR Time Horizons—65.4%

Reasoning o3 leads

Mistral Large: 15.8 (#310), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkMistral Largeo3
SimpleBench22.5%53.1%
CritPt0%1.4%
LMArena Hard Prompts12571402
DTBench65.1%84.8%
LMCA16.7%39.7%
Epoch Capabilities Index128.52146.86
ForecastBench57.162.5
ARC-AGI-2—6.5%
Kagi LLM Benchmark—67.6%
ARC-AGI-1—60.8%
Chess Puzzles—38%
EnigmaEval—13.1%
LiveBench Reasoning43.5%—
Mystery Game Puzzles—29%
LiveBench Data Analysis50.1%—
LiveBench48.4%—

Math o3 leads

Mistral Large: 18.2 (#291), o3: 50.2 (#58)

Math benchmarks
BenchmarkMistral Largeo3
OTIS Mock AIME 2024-20258.5%84.4%
Omni-MATH28.1%71.4%
LMArena Math12621426
MATH Level 550.3%97.8%
FrontierMath (Feb 2025 set)0.3%18.7%
FrontierMath (Tiers 1-3)—33.3%
LiveBench Math42.5%—
FrontierMath Tier 4 (v1)—2.1%

Knowledge o3 leads

Mistral Large: 30.1 (#230), o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkMistral Largeo3
GPQA Diamond51.3%81.8%
MMLU-Pro59.9%85.9%
Confabulations21.4%14.4%
GPQA (HELM)43.5%75.3%
LMArena Expert12321402
Humanity's Last Exam—20.3%
SimpleQA Verified—49.4%
Vectara Hallucination Rate4.5%—
MMLU80%—

Multimodal Not comparable

Mistral Large: —, o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkMistral Largeo3
LMArena Vision—1214
GeoBench—74%
VPCT—52%

Multilingual o3 leads

Mistral Large: 40.0 (#219), o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkMistral Largeo3
LMArena Non-English12371401
LMArena Chinese12401437
LMArena French13251430
LMArena German12541420
LMArena Japanese11881403
LMArena Korean12021370
LMArena Russian12571406
LMArena Spanish12681395

Instruction Following o3 leads

Mistral Large: 67.9 (#191), o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkMistral Largeo3
IFEval87.7%86.9%
LMArena Instruction Following12491368
LiveBench Instruction Following67.9%—

Long Context o3 leads

Mistral Large: 38.3 (#199), o3: 53.3 (#6)

Long Context benchmarks
BenchmarkMistral Largeo3
LMArena Longer Query12611372
Fiction.LiveBench—88.9%
CL-bench—17.8%

Writing & Preference o3 leads

Mistral Large: 40.7 (#242), o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkMistral Largeo3
LMArena Text12661410
LMArena Creative Writing12431359
Short-Story Creative Writing69%83.9%
EQ-Bench Creative Writing9851676
WildBench80.1%86.1%
LMArena Multi-Turn12601405
LiveBench Language39.4%—

Frequently asked questions

Is Mistral Large better than o3?

o3 is the stronger model overall, scoring 47.5 to 31.9 on the Noometry Index.

Which is cheaper, Mistral Large or o3?

Mistral Large is cheaper. It lists at $2 per million input tokens and $6 per million output tokens; o3 lists at $2 and $8.

Is Mistral Large or o3 better for coding?

o3 scores higher on coding benchmarks: 46.8 versus 34.3 in the Noometry coding category.

Which has the bigger context window?

o3 does, with 200K tokens against 131K.

How many benchmarks do Mistral Large and o3 share?

37 benchmarks have published results for both models. Mistral Large has 51 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper