Model comparison

Mistral Medium vs Pixtral Large

Mistral Medium is the stronger model overall, scoring 36.3 to 32.2 on the Noometry Index.

Last verified . 1 shared benchmarks.

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

Pixtral Large Mistral AI

32.2

Rank #259 Reported

Summary

  • They share 1 benchmark with published results for both. Mistral Medium scores higher in 3 categories and Pixtral Large in 0 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Medium leads 60.0 to 32.9.
  • Both cost about the same: $1.50 input and $7.50 output per million tokens.
  • Mistral Medium accepts more context: 262K tokens versus 128K.

Side by side

Mistral Medium and Pixtral Large specifications
Mistral MediumPixtral Large
ProviderMistral AIMistral AI
Noometry Index36.332.2
Released2023-12-112024-11-01
WeightsOpenOpen
Context window262K128K
Max output262K128K
Input $ / M tokens$1.50$2
Output $ / M tokens$7.50$6
Results tracked363

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Mistral Medium: 34.2 (#243), Pixtral Large: —

Coding benchmarks
BenchmarkMistral MediumPixtral Large
FrontierCode8%—
SciCode40.2%—
WeirdML43.7%—
LMArena Coding1434—
ALE-Bench763.98—

Agentic & Tool Use Not comparable

Mistral Medium: 28.3 (#90), Pixtral Large: —

Agentic & Tool Use benchmarks
BenchmarkMistral MediumPixtral Large
Berkeley Function Calling Leaderboard37.7%—

Reasoning Mistral Medium leads

Mistral Medium: 24.0 (#167), Pixtral Large: 21.7 (#218)

Reasoning benchmarks
BenchmarkMistral MediumPixtral Large
Kagi LLM Benchmark50%—
CritPt0%—
EnigmaEval—0.8%
LMArena Hard Prompts1426—
DTBench75.5%—
LMCA26.1%—
Surface Evolver Bench26.9%—

Math Not comparable

Mistral Medium: 28.1 (#245), Pixtral Large: —

Math benchmarks
BenchmarkMistral MediumPixtral Large
OTIS Mock AIME 2024-202532.2%—
ProofBench9%—
LMArena Math1408—
MATH Level 581.6%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Not comparable

Mistral Medium: 25.0 (#265), Pixtral Large: —

Knowledge benchmarks
BenchmarkMistral MediumPixtral Large
GPQA Diamond59.5%—
Humanity's Last Exam4.5%—
Vectara Hallucination Rate22.7%—
LMArena Expert1408—

Multimodal Mistral Medium leads

Mistral Medium: 35.3 (#88), Pixtral Large: 30.6 (#111)

Multimodal benchmarks
BenchmarkMistral MediumPixtral Large
LMArena Vision11721089

Multilingual Not comparable

Mistral Medium: 52.1 (#91), Pixtral Large: —

Multilingual benchmarks
BenchmarkMistral MediumPixtral Large
LMArena Non-English1408—
LMArena Chinese1447—
LMArena French1459—
LMArena German1432—
LMArena Japanese1378—
LMArena Korean1380—
LMArena Russian1411—
LMArena Spanish1433—

Instruction Following Not comparable

Mistral Medium: 73.7 (#116), Pixtral Large: —

Instruction Following benchmarks
BenchmarkMistral MediumPixtral Large
LMArena Instruction Following1398—

Long Context Not comparable

Mistral Medium: 42.9 (#114), Pixtral Large: —

Long Context benchmarks
BenchmarkMistral MediumPixtral Large
LMArena Longer Query1406—

Writing & Preference Mistral Medium leads

Mistral Medium: 60.0 (#103), Pixtral Large: 32.9 (#278)

Writing & Preference benchmarks
BenchmarkMistral MediumPixtral Large
LMArena Text1424—
LMArena Creative Writing1391—
Short-Story Creative Writing77.3%—
EQ-Bench Creative Writing—988
LMArena Multi-Turn1418—

Frequently asked questions

Is Mistral Medium better than Pixtral Large?

Mistral Medium is the stronger model overall, scoring 36.3 to 32.2 on the Noometry Index.

Which is cheaper, Mistral Medium or Pixtral Large?

Pixtral Large is cheaper. It lists at $2 per million input tokens and $6 per million output tokens; Mistral Medium lists at $1.50 and $7.50.

Which has the bigger context window?

Mistral Medium does, with 262K tokens against 128K.

How many benchmarks do Mistral Medium and Pixtral Large share?

1 benchmark has published results for both models. Mistral Medium has 36 scored results on Noometry and Pixtral Large has 3.

Related comparisons

Go deeper