Model comparison

Olmo 3.1 32b Instruct vs Qwen Max

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 34.7 on the Noometry Index.

Last verified . 16 shared benchmarks.

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Olmo 3.1 32b Instruct scores higher in 8 categories and Qwen Max in 0 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where Olmo 3.1 32b Instruct leads 36.3 to 22.3.
  • Olmo 3.1 32b Instruct has downloadable open weights; the other is API-only.

Side by side

Olmo 3.1 32b Instruct and Qwen Max specifications
Olmo 3.1 32b InstructQwen Max
ProviderAllen Institute for AI (Ai2)Alibaba (Qwen)
Noometry Index39.434.7
Released—2024-04-03
WeightsOpenProprietary
Context window—33K
Max output—8K
Input $ / M tokens—$1.60
Output $ / M tokens—$6.40
Results tracked1623

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 39.5 (#157), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkOlmo 3.1 32b InstructQwen Max
LMArena Coding13471288
Aider Polyglot—21.8%

Reasoning Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 26.4 (#132), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkOlmo 3.1 32b InstructQwen Max
LMArena Hard Prompts13221269

Math Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 36.3 (#167), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkOlmo 3.1 32b InstructQwen Max
LMArena Math13051275
OTIS Mock AIME 2024-2025—16.1%
MATH Level 5—67.2%
FrontierMath (Feb 2025 set)—1%

Knowledge Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 36.1 (#175), Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkOlmo 3.1 32b InstructQwen Max
LMArena Expert13081248
GPQA Diamond—56.1%

Multilingual Too close to call

Olmo 3.1 32b Instruct: 42.6 (#191), Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkOlmo 3.1 32b InstructQwen Max
LMArena Non-English12751263
LMArena Chinese13041254
LMArena French13281330
LMArena German12821254
LMArena Korean12061142
LMArena Russian12681274
LMArena Spanish13361290
LMArena Japanese—1205

Instruction Following Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 68.6 (#187), Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b InstructQwen Max
LMArena Instruction Following12991262

Long Context Too close to call

Olmo 3.1 32b Instruct: 39.9 (#166), Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkOlmo 3.1 32b InstructQwen Max
LMArena Longer Query13121288
Fiction.LiveBench—66.7%

Writing & Preference Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 50.2 (#185), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b InstructQwen Max
LMArena Text13111282
LMArena Creative Writing12641248
LMArena Multi-Turn13091277

Frequently asked questions

Is Olmo 3.1 32b Instruct better than Qwen Max?

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 34.7 on the Noometry Index.

Is Olmo 3.1 32b Instruct or Qwen Max better for coding?

Olmo 3.1 32b Instruct scores higher on coding benchmarks: 39.5 versus 30.7 in the Noometry coding category.

How many benchmarks do Olmo 3.1 32b Instruct and Qwen Max share?

16 benchmarks have published results for both models. Olmo 3.1 32b Instruct has 16 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper