Model comparison

Mistral Small 3.1 vs Pixtral Large

Mistral Small 3.1 and Pixtral Large score almost the same on the Noometry Index (31.7 vs 32.2), so choose on price, context window or the category you care about most.

Last verified . 2 shared benchmarks.

Mistral Small 3.1 Mistral AI

31.7

Rank #269 Confirmed

Pixtral Large Mistral AI

32.2

Rank #259 Reported

Summary

  • They share 2 benchmarks with published results for both. Mistral Small 3.1 scores higher in 2 categories and Pixtral Large in 1 category; 3 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Small 3.1 leads 37.0 to 32.9.
  • Mistral Small 3.1 is cheaper at $0.35 / $0.56 per million input/output tokens, against $2 / $6 for Pixtral Large.

Side by side

Mistral Small 3.1 and Pixtral Large specifications
Mistral Small 3.1Pixtral Large
ProviderMistral AIMistral AI
Noometry Index31.732.2
Released2025-03-172024-11-01
WeightsOpenOpen
Context window128K128K
Max output102K128K
Input $ / M tokens$0.35$2
Output $ / M tokens$0.56$6
Results tracked283

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Mistral Small 3.1: 38.3 (#179), Pixtral Large: —

Coding benchmarks
BenchmarkMistral Small 3.1Pixtral Large
LMArena Coding1309—

Reasoning Pixtral Large leads

Mistral Small 3.1: 19.7 (#254), Pixtral Large: 21.7 (#218)

Reasoning benchmarks
BenchmarkMistral Small 3.1Pixtral Large
Chess Puzzles1%—
EnigmaEval—0.8%
LMArena Hard Prompts1278—
Epoch Capabilities Index127.48—

Math Not comparable

Mistral Small 3.1: 14.7 (#301), Pixtral Large: —

Math benchmarks
BenchmarkMistral Small 3.1Pixtral Large
OTIS Mock AIME 2024-20253.9%—
Omni-MATH24.8%—
LMArena Math1262—

Knowledge Not comparable

Mistral Small 3.1: 22.6 (#271), Pixtral Large: —

Knowledge benchmarks
BenchmarkMistral Small 3.1Pixtral Large
GPQA Diamond41.9%—
MMLU-Pro61%—
GPQA (HELM)39.2%—
LMArena Expert1257—

Multimodal Mistral Small 3.1 leads

Mistral Small 3.1: 33.2 (#99), Pixtral Large: 30.6 (#111)

Multimodal benchmarks
BenchmarkMistral Small 3.1Pixtral Large
LMArena Vision11361089

Multilingual Not comparable

Mistral Small 3.1: 41.2 (#209), Pixtral Large: —

Multilingual benchmarks
BenchmarkMistral Small 3.1Pixtral Large
LMArena Non-English1255—
LMArena Chinese1253—
LMArena French1273—
LMArena German1266—
LMArena Japanese1208—
LMArena Korean1206—
LMArena Russian1263—
LMArena Spanish1283—

Instruction Following Not comparable

Mistral Small 3.1: 63.6 (#230), Pixtral Large: —

Instruction Following benchmarks
BenchmarkMistral Small 3.1Pixtral Large
IFEval75%—
LMArena Instruction Following1264—

Long Context Not comparable

Mistral Small 3.1: 39.5 (#178), Pixtral Large: —

Long Context benchmarks
BenchmarkMistral Small 3.1Pixtral Large
LMArena Longer Query1299—

Writing & Preference Mistral Small 3.1 leads

Mistral Small 3.1: 37.0 (#259), Pixtral Large: 32.9 (#278)

Writing & Preference benchmarks
BenchmarkMistral Small 3.1Pixtral Large
EQ-Bench Creative Writing761988
LMArena Text1277—
LMArena Creative Writing1253—
WildBench78.8%—
LMArena Multi-Turn1270—

Frequently asked questions

Is Mistral Small 3.1 better than Pixtral Large?

Mistral Small 3.1 and Pixtral Large score almost the same on the Noometry Index (31.7 vs 32.2), so choose on price, context window or the category you care about most.

Which is cheaper, Mistral Small 3.1 or Pixtral Large?

Mistral Small 3.1 is cheaper. It lists at $0.35 per million input tokens and $0.56 per million output tokens; Pixtral Large lists at $2 and $6.

Which has the bigger context window?

Both accept 128K tokens.

How many benchmarks do Mistral Small 3.1 and Pixtral Large share?

2 benchmarks have published results for both models. Mistral Small 3.1 has 28 scored results on Noometry and Pixtral Large has 3.

Related comparisons

Go deeper