Model comparison

Olmo 3.1 32b Think vs Qwen Max

Olmo 3.1 32b Think is the stronger model overall, scoring 37.9 to 34.7 on the Noometry Index.

Last verified . 15 shared benchmarks.

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Olmo 3.1 32b Think scores higher in 4 categories and Qwen Max in 4 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in math, where Olmo 3.1 32b Think leads 36.3 to 22.3.
  • Olmo 3.1 32b Think has downloadable open weights; the other is API-only.

Side by side

Olmo 3.1 32b Think and Qwen Max specifications
Olmo 3.1 32b ThinkQwen Max
ProviderAllen Institute for AI (Ai2)Alibaba (Qwen)
Noometry Index37.934.7
Released—2024-04-03
WeightsOpenProprietary
Context window—33K
Max output—8K
Input $ / M tokens—$1.60
Output $ / M tokens—$6.40
Results tracked1523

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Think leads

Olmo 3.1 32b Think: 37.7 (#189), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen Max
LMArena Coding12911288
Aider Polyglot—21.8%

Reasoning Too close to call

Olmo 3.1 32b Think: 25.2 (#150), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen Max
LMArena Hard Prompts12721269

Math Olmo 3.1 32b Think leads

Olmo 3.1 32b Think: 36.3 (#168), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen Max
LMArena Math13051275
OTIS Mock AIME 2024-2025—16.1%
MATH Level 5—67.2%
FrontierMath (Feb 2025 set)—1%

Knowledge Olmo 3.1 32b Think leads

Olmo 3.1 32b Think: 35.7 (#181), Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen Max
LMArena Expert12951248
GPQA Diamond—56.1%

Multilingual Qwen Max leads

Olmo 3.1 32b Think: 38.1 (#231), Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen Max
LMArena Non-English12091263
LMArena Chinese12421254
LMArena French12601330
LMArena German12621254
LMArena Russian11931274
LMArena Spanish12891290
LMArena Japanese—1205
LMArena Korean—1142

Instruction Following Too close to call

Olmo 3.1 32b Think: 65.6 (#218), Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen Max
LMArena Instruction Following12471262

Long Context Too close to call

Olmo 3.1 32b Think: 38.6 (#195), Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen Max
LMArena Longer Query12721288
Fiction.LiveBench—66.7%

Writing & Preference Qwen Max leads

Olmo 3.1 32b Think: 46.2 (#220), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen Max
LMArena Text12721282
LMArena Creative Writing12261248
LMArena Multi-Turn12521277

Frequently asked questions

Is Olmo 3.1 32b Think better than Qwen Max?

Olmo 3.1 32b Think is the stronger model overall, scoring 37.9 to 34.7 on the Noometry Index.

Is Olmo 3.1 32b Think or Qwen Max better for coding?

Olmo 3.1 32b Think scores higher on coding benchmarks: 37.7 versus 30.7 in the Noometry coding category.

How many benchmarks do Olmo 3.1 32b Think and Qwen Max share?

15 benchmarks have published results for both models. Olmo 3.1 32b Think has 15 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper