Model comparison

Olmo 3.1 32b Think vs Qwen3.5 Max Preview

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 37.9 on the Noometry Index.

Last verified . 15 shared benchmarks.

Summary

  • They share 15 benchmarks with published results for both. Olmo 3.1 32b Think scores higher in 0 categories and Qwen3.5 Max Preview in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen3.5 Max Preview leads 66.0 to 46.2.
  • Olmo 3.1 32b Think has downloadable open weights; the other is API-only.

Side by side

Olmo 3.1 32b Think and Qwen3.5 Max Preview specifications
Olmo 3.1 32b ThinkQwen3.5 Max Preview
ProviderAllen Institute for AI (Ai2)Alibaba (Qwen)
Noometry Index37.945.3
Released——
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1517

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.5 Max Preview leads

Olmo 3.1 32b Think: 37.7 (#189), Qwen3.5 Max Preview: 44.0 (#77)

Coding benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen3.5 Max Preview
LMArena Coding12911487

Reasoning Qwen3.5 Max Preview leads

Olmo 3.1 32b Think: 25.2 (#150), Qwen3.5 Max Preview: 30.8 (#84)

Reasoning benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen3.5 Max Preview
LMArena Hard Prompts12721483

Math Qwen3.5 Max Preview leads

Olmo 3.1 32b Think: 36.3 (#168), Qwen3.5 Max Preview: 40.1 (#94)

Math benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen3.5 Max Preview
LMArena Math13051474

Knowledge Qwen3.5 Max Preview leads

Olmo 3.1 32b Think: 35.7 (#181), Qwen3.5 Max Preview: 41.8 (#107)

Knowledge benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen3.5 Max Preview
LMArena Expert12951489

Multilingual Qwen3.5 Max Preview leads

Olmo 3.1 32b Think: 38.1 (#231), Qwen3.5 Max Preview: 56.2 (#22)

Multilingual benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen3.5 Max Preview
LMArena Non-English12091465
LMArena Chinese12421534
LMArena French12601484
LMArena German12621487
LMArena Russian11931471
LMArena Spanish12891470
LMArena Japanese—1495
LMArena Korean—1438

Instruction Following Qwen3.5 Max Preview leads

Olmo 3.1 32b Think: 65.6 (#218), Qwen3.5 Max Preview: 77.0 (#31)

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen3.5 Max Preview
LMArena Instruction Following12471467

Long Context Qwen3.5 Max Preview leads

Olmo 3.1 32b Think: 38.6 (#195), Qwen3.5 Max Preview: 45.2 (#45)

Long Context benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen3.5 Max Preview
LMArena Longer Query12721476

Writing & Preference Qwen3.5 Max Preview leads

Olmo 3.1 32b Think: 46.2 (#220), Qwen3.5 Max Preview: 66.0 (#41)

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen3.5 Max Preview
LMArena Text12721470
LMArena Creative Writing12261464
LMArena Multi-Turn12521478

Frequently asked questions

Is Olmo 3.1 32b Think better than Qwen3.5 Max Preview?

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 37.9 on the Noometry Index.

Is Olmo 3.1 32b Think or Qwen3.5 Max Preview better for coding?

Qwen3.5 Max Preview scores higher on coding benchmarks: 44.0 versus 37.7 in the Noometry coding category.

How many benchmarks do Olmo 3.1 32b Think and Qwen3.5 Max Preview share?

15 benchmarks have published results for both models. Olmo 3.1 32b Think has 15 scored results on Noometry and Qwen3.5 Max Preview has 17.

Related comparisons

Go deeper