Model comparison

Olmo 3.1 32b Instruct vs QwQ-32B

Olmo 3.1 32b Instruct and QwQ-32B score almost the same on the Noometry Index (39.4 vs 39.8), so choose on price, context window or the category you care about most.

Last verified . 16 shared benchmarks.

QwQ-32B Alibaba (Qwen)

39.8

Rank #159 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Olmo 3.1 32b Instruct scores higher in 2 categories and QwQ-32B in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in long context, where QwQ-32B leads 49.0 to 39.9.

Side by side

Olmo 3.1 32b Instruct and QwQ-32B specifications
Olmo 3.1 32b InstructQwQ-32B
ProviderAllen Institute for AI (Ai2)Alibaba (Qwen)
Noometry Index39.439.8
Released—2024-11-28
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1636

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 39.5 (#157), QwQ-32B: 35.4 (#226)

Coding benchmarks
BenchmarkOlmo 3.1 32b InstructQwQ-32B
LMArena Coding13471333
Aider Polyglot—20.9%
BigCodeBench Instruct—44.6%
LiveBench Coding—72.2%
BigCodeBench Complete—54.4%

Reasoning Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 26.4 (#132), QwQ-32B: 23.7 (#174)

Reasoning benchmarks
BenchmarkOlmo 3.1 32b InstructQwQ-32B
LMArena Hard Prompts13221325
Chess Puzzles—5%
LiveBench Reasoning—83.5%
LiveBench Data Analysis—65%
Epoch Capabilities Index—137.6
ForecastBench—58.3
LiveBench—72%

Math QwQ-32B leads

Olmo 3.1 32b Instruct: 36.3 (#167), QwQ-32B: 38.0 (#143)

Math benchmarks
BenchmarkOlmo 3.1 32b InstructQwQ-32B
LMArena Math13051359
OTIS Mock AIME 2024-2025—59.2%
LiveBench Math—77.8%

Knowledge QwQ-32B leads

Olmo 3.1 32b Instruct: 36.1 (#175), QwQ-32B: 37.2 (#158)

Knowledge benchmarks
BenchmarkOlmo 3.1 32b InstructQwQ-32B
LMArena Expert13081324
GPQA Diamond—65.3%
Confabulations—15.6%

Multilingual QwQ-32B leads

Olmo 3.1 32b Instruct: 42.6 (#191), QwQ-32B: 44.8 (#176)

Multilingual benchmarks
BenchmarkOlmo 3.1 32b InstructQwQ-32B
LMArena Non-English12751305
LMArena Chinese13041378
LMArena French13281336
LMArena German12821313
LMArena Korean12061279
LMArena Russian12681297
LMArena Spanish13361354
LMArena Japanese—1262

Instruction Following QwQ-32B leads

Olmo 3.1 32b Instruct: 68.6 (#187), QwQ-32B: 72.6 (#137)

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b InstructQwQ-32B
LMArena Instruction Following12991297
LiveBench Instruction Following—81.8%

Long Context QwQ-32B leads

Olmo 3.1 32b Instruct: 39.9 (#166), QwQ-32B: 49.0 (#11)

Long Context benchmarks
BenchmarkOlmo 3.1 32b InstructQwQ-32B
LMArena Longer Query13121308
Fiction.LiveBench—83.3%

Writing & Preference Too close to call

Olmo 3.1 32b Instruct: 50.2 (#185), QwQ-32B: 50.6 (#180)

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b InstructQwQ-32B
LMArena Text13111329
LMArena Creative Writing12641288
LMArena Multi-Turn13091314
Short-Story Creative Writing—80.2%
EQ-Bench Creative Writing—1257
LiveBench Language—51.4%

Frequently asked questions

Is Olmo 3.1 32b Instruct better than QwQ-32B?

Olmo 3.1 32b Instruct and QwQ-32B score almost the same on the Noometry Index (39.4 vs 39.8), so choose on price, context window or the category you care about most.

Is Olmo 3.1 32b Instruct or QwQ-32B better for coding?

Olmo 3.1 32b Instruct scores higher on coding benchmarks: 39.5 versus 35.4 in the Noometry coding category.

How many benchmarks do Olmo 3.1 32b Instruct and QwQ-32B share?

16 benchmarks have published results for both models. Olmo 3.1 32b Instruct has 16 scored results on Noometry and QwQ-32B has 36.

Related comparisons

Go deeper