Model comparison

Mistral Large vs Mistral Medium

Mistral Medium is the stronger model overall, scoring 36.3 to 31.9 on the Noometry Index.

Last verified . 29 shared benchmarks.

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

Summary

  • They share 29 benchmarks with published results for both. Mistral Large scores higher in 3 categories and Mistral Medium in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Medium leads 60.0 to 40.7.
  • The biggest single-benchmark swing is MATH Level 5: 50.3% for Mistral Large and 81.6% for Mistral Medium.
  • Both cost about the same: $2 input and $6 output per million tokens.
  • Mistral Medium accepts more context: 262K tokens versus 131K.

Side by side

Mistral Large and Mistral Medium specifications
Mistral LargeMistral Medium
ProviderMistral AIMistral AI
Noometry Index31.936.3
Released2024-02-262023-12-11
WeightsOpenOpen
Context window131K262K
Max output16K262K
Input $ / M tokens$2$1.50
Output $ / M tokens$6$7.50
Results tracked5136

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Mistral Large: 34.3 (#240), Mistral Medium: 34.2 (#243)

Coding benchmarks
BenchmarkMistral LargeMistral Medium
SciCode36.2%40.2%
LMArena Coding12771434
ALE-Bench264.7763.98
FrontierCode—8%
WeirdML—43.7%
BigCodeBench Instruct30%—
LiveBench Coding47.1%—
BigCodeBench Complete38.3%—
HumanEval+62.2%—
MBPP+59.5%—

Agentic & Tool Use Too close to call

Mistral Large: 28.6 (#89), Mistral Medium: 28.3 (#90)

Agentic & Tool Use benchmarks
BenchmarkMistral LargeMistral Medium
Berkeley Function Calling Leaderboard38.4%37.7%

Reasoning Mistral Medium leads

Mistral Large: 15.8 (#310), Mistral Medium: 24.0 (#167)

Reasoning benchmarks
BenchmarkMistral LargeMistral Medium
CritPt0%0%
LMArena Hard Prompts12571426
DTBench65.1%75.5%
LMCA16.7%26.1%
SimpleBench22.5%—
Kagi LLM Benchmark—50%
LiveBench Reasoning43.5%—
LiveBench Data Analysis50.1%—
Surface Evolver Bench—26.9%
Epoch Capabilities Index128.52—
ForecastBench57.1—
LiveBench48.4%—

Math Mistral Medium leads

Mistral Large: 18.2 (#291), Mistral Medium: 28.1 (#245)

Math benchmarks
BenchmarkMistral LargeMistral Medium
OTIS Mock AIME 2024-20258.5%32.2%
LMArena Math12621408
MATH Level 550.3%81.6%
FrontierMath (Feb 2025 set)0.3%0.3%
ProofBench—9%
Omni-MATH28.1%—
LiveBench Math42.5%—

Knowledge Mistral Large leads

Mistral Large: 30.1 (#230), Mistral Medium: 25.0 (#265)

Knowledge benchmarks
BenchmarkMistral LargeMistral Medium
GPQA Diamond51.3%59.5%
Vectara Hallucination Rate4.5%22.7%
LMArena Expert12321408
Humanity's Last Exam—4.5%
MMLU-Pro59.9%—
Confabulations21.4%—
GPQA (HELM)43.5%—
MMLU80%—

Multimodal Not comparable

Mistral Large: —, Mistral Medium: 35.3 (#88)

Multimodal benchmarks
BenchmarkMistral LargeMistral Medium
LMArena Vision—1172

Multilingual Mistral Medium leads

Mistral Large: 40.0 (#219), Mistral Medium: 52.1 (#91)

Multilingual benchmarks
BenchmarkMistral LargeMistral Medium
LMArena Non-English12371408
LMArena Chinese12401447
LMArena French13251459
LMArena German12541432
LMArena Japanese11881378
LMArena Korean12021380
LMArena Russian12571411
LMArena Spanish12681433

Instruction Following Mistral Medium leads

Mistral Large: 67.9 (#191), Mistral Medium: 73.7 (#116)

Instruction Following benchmarks
BenchmarkMistral LargeMistral Medium
LMArena Instruction Following12491398
LiveBench Instruction Following67.9%—
IFEval87.7%—

Long Context Mistral Medium leads

Mistral Large: 38.3 (#199), Mistral Medium: 42.9 (#114)

Long Context benchmarks
BenchmarkMistral LargeMistral Medium
LMArena Longer Query12611406

Writing & Preference Mistral Medium leads

Mistral Large: 40.7 (#242), Mistral Medium: 60.0 (#103)

Writing & Preference benchmarks
BenchmarkMistral LargeMistral Medium
LMArena Text12661424
LMArena Creative Writing12431391
Short-Story Creative Writing69%77.3%
LMArena Multi-Turn12601418
EQ-Bench Creative Writing985—
WildBench80.1%—
LiveBench Language39.4%—

Frequently asked questions

Is Mistral Large better than Mistral Medium?

Mistral Medium is the stronger model overall, scoring 36.3 to 31.9 on the Noometry Index.

Which is cheaper, Mistral Large or Mistral Medium?

Mistral Medium is cheaper. It lists at $1.50 per million input tokens and $7.50 per million output tokens; Mistral Large lists at $2 and $6.

Is Mistral Large or Mistral Medium better for coding?

They score almost the same on coding (34.3 vs 34.2); test both on your own repository before choosing.

Which has the bigger context window?

Mistral Medium does, with 262K tokens against 131K.

How many benchmarks do Mistral Large and Mistral Medium share?

29 benchmarks have published results for both models. Mistral Large has 51 scored results on Noometry and Mistral Medium has 36.

Related comparisons

Go deeper