Model comparison

Olmo 3.1 32b Instruct vs Pixtral Large

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 32.2 on the Noometry Index.

Last verified . 0 shared benchmarks.

Summary

  • The widest gap is in writing & preference, where Olmo 3.1 32b Instruct leads 50.2 to 32.9.

Side by side

Olmo 3.1 32b Instruct and Pixtral Large specifications
Olmo 3.1 32b InstructPixtral Large
ProviderAllen Institute for AI (Ai2)Mistral AI
Noometry Index39.432.2
Released—2024-11-01
WeightsOpenOpen
Context window—128K
Max output—128K
Input $ / M tokens—$2
Output $ / M tokens—$6
Results tracked163

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Olmo 3.1 32b Instruct: 39.5 (#157), Pixtral Large: —

Coding benchmarks
BenchmarkOlmo 3.1 32b InstructPixtral Large
LMArena Coding1347—

Reasoning Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 26.4 (#132), Pixtral Large: 21.7 (#218)

Reasoning benchmarks
BenchmarkOlmo 3.1 32b InstructPixtral Large
EnigmaEval—0.8%
LMArena Hard Prompts1322—

Math Not comparable

Olmo 3.1 32b Instruct: 36.3 (#167), Pixtral Large: —

Math benchmarks
BenchmarkOlmo 3.1 32b InstructPixtral Large
LMArena Math1305—

Knowledge Not comparable

Olmo 3.1 32b Instruct: 36.1 (#175), Pixtral Large: —

Knowledge benchmarks
BenchmarkOlmo 3.1 32b InstructPixtral Large
LMArena Expert1308—

Multimodal Not comparable

Olmo 3.1 32b Instruct: —, Pixtral Large: 30.6 (#111)

Multimodal benchmarks
BenchmarkOlmo 3.1 32b InstructPixtral Large
LMArena Vision—1089

Multilingual Not comparable

Olmo 3.1 32b Instruct: 42.6 (#191), Pixtral Large: —

Multilingual benchmarks
BenchmarkOlmo 3.1 32b InstructPixtral Large
LMArena Non-English1275—
LMArena Chinese1304—
LMArena French1328—
LMArena German1282—
LMArena Korean1206—
LMArena Russian1268—
LMArena Spanish1336—

Instruction Following Not comparable

Olmo 3.1 32b Instruct: 68.6 (#187), Pixtral Large: —

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b InstructPixtral Large
LMArena Instruction Following1299—

Long Context Not comparable

Olmo 3.1 32b Instruct: 39.9 (#166), Pixtral Large: —

Long Context benchmarks
BenchmarkOlmo 3.1 32b InstructPixtral Large
LMArena Longer Query1312—

Writing & Preference Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 50.2 (#185), Pixtral Large: 32.9 (#278)

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b InstructPixtral Large
LMArena Text1311—
LMArena Creative Writing1264—
EQ-Bench Creative Writing—988
LMArena Multi-Turn1309—

Frequently asked questions

Is Olmo 3.1 32b Instruct better than Pixtral Large?

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 32.2 on the Noometry Index.

How many benchmarks do Olmo 3.1 32b Instruct and Pixtral Large share?

0 benchmarks have published results for both models. Olmo 3.1 32b Instruct has 16 scored results on Noometry and Pixtral Large has 3.

Related comparisons

Go deeper