Model comparison

Mistral Small vs Pixtral Large

Mistral Small is the stronger model overall, scoring 33.4 to 32.2 on the Noometry Index.

Last verified . 1 shared benchmarks.

Mistral Small Mistral AI

33.4

Rank #243 Confirmed

Pixtral Large Mistral AI

32.2

Rank #259 Reported

Summary

  • They share 1 benchmark with published results for both. Mistral Small scores higher in 2 categories and Pixtral Large in 1 category; 3 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Small leads 52.5 to 32.9.
  • Mistral Small is cheaper at $0.15 / $0.60 per million input/output tokens, against $2 / $6 for Pixtral Large.
  • Mistral Small accepts more context: 262K tokens versus 128K.

Side by side

Mistral Small and Pixtral Large specifications
Mistral SmallPixtral Large
ProviderMistral AIMistral AI
Noometry Index33.432.2
Released2024-02-262024-11-01
WeightsOpenOpen
Context window262K128K
Max output256K128K
Input $ / M tokens$0.15$2
Output $ / M tokens$0.60$6
Results tracked393

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Mistral Small: 34.0 (#247), Pixtral Large: —

Coding benchmarks
BenchmarkMistral SmallPixtral Large
SciCode26.5%—
BigCodeBench Instruct36.1%—
LiveBench Coding36.2%—
LMArena Coding1362—
BigCodeBench Complete46.6%—
ALE-Bench497.62—

Agentic & Tool Use Not comparable

Mistral Small: 28.1 (#93), Pixtral Large: —

Agentic & Tool Use benchmarks
BenchmarkMistral SmallPixtral Large
Berkeley Function Calling Leaderboard37.1%—

Reasoning Pixtral Large leads

Mistral Small: 19.8 (#250), Pixtral Large: 21.7 (#218)

Reasoning benchmarks
BenchmarkMistral SmallPixtral Large
Kagi LLM Benchmark37.8%—
CritPt0%—
EnigmaEval—0.8%
LiveBench Reasoning44.8%—
LMArena Hard Prompts1335—
DTBench70.9%—
LiveBench Data Analysis53.7%—
LMCA20.6%—
LiveBench44%—

Math Not comparable

Mistral Small: 16.4 (#293), Pixtral Large: —

Math benchmarks
BenchmarkMistral SmallPixtral Large
OTIS Mock AIME 2024-20255.8%—
LiveBench Math39.9%—
LMArena Math1341—
MATH Level 546.8%—

Knowledge Not comparable

Mistral Small: 31.0 (#222), Pixtral Large: —

Knowledge benchmarks
BenchmarkMistral SmallPixtral Large
GPQA Diamond47.5%—
Vectara Hallucination Rate5.1%—
LMArena Expert1291—
MMLU68.7%—

Multimodal Mistral Small leads

Mistral Small: 33.5 (#96), Pixtral Large: 30.6 (#111)

Multimodal benchmarks
BenchmarkMistral SmallPixtral Large
LMArena Vision11421089

Multilingual Not comparable

Mistral Small: 45.5 (#169), Pixtral Large: —

Multilingual benchmarks
BenchmarkMistral SmallPixtral Large
LMArena Non-English1315—
LMArena Chinese1340—
LMArena French1337—
LMArena German1340—
LMArena Japanese1275—
LMArena Korean1259—
LMArena Russian1324—
LMArena Spanish1346—

Instruction Following Not comparable

Mistral Small: 66.4 (#209), Pixtral Large: —

Instruction Following benchmarks
BenchmarkMistral SmallPixtral Large
LiveBench Instruction Following63.7%—
LMArena Instruction Following1310—

Long Context Not comparable

Mistral Small: 40.4 (#156), Pixtral Large: —

Long Context benchmarks
BenchmarkMistral SmallPixtral Large
LMArena Longer Query1327—

Writing & Preference Mistral Small leads

Mistral Small: 52.5 (#171), Pixtral Large: 32.9 (#278)

Writing & Preference benchmarks
BenchmarkMistral SmallPixtral Large
LMArena Text1338—
LMArena Creative Writing1305—
EQ-Bench Creative Writing—988
LMArena Multi-Turn1344—
LiveBench Language30.5%—

Frequently asked questions

Is Mistral Small better than Pixtral Large?

Mistral Small is the stronger model overall, scoring 33.4 to 32.2 on the Noometry Index.

Which is cheaper, Mistral Small or Pixtral Large?

Mistral Small is cheaper. It lists at $0.15 per million input tokens and $0.60 per million output tokens; Pixtral Large lists at $2 and $6.

Which has the bigger context window?

Mistral Small does, with 262K tokens against 128K.

How many benchmarks do Mistral Small and Pixtral Large share?

1 benchmark has published results for both models. Mistral Small has 39 scored results on Noometry and Pixtral Large has 3.

Related comparisons

Go deeper