Model comparison

Olmo 3.1 32b Instruct vs Qwen Plus

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 37.1 on the Noometry Index.

Last verified . 12 shared benchmarks.

Qwen Plus Alibaba (Qwen)

37.1

Rank #210 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Olmo 3.1 32b Instruct scores higher in 3 categories and Qwen Plus in 5 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in math, where Olmo 3.1 32b Instruct leads 36.3 to 23.3.
  • Olmo 3.1 32b Instruct has downloadable open weights; the other is API-only.

Side by side

Olmo 3.1 32b Instruct and Qwen Plus specifications
Olmo 3.1 32b InstructQwen Plus
ProviderAllen Institute for AI (Ai2)Alibaba (Qwen)
Noometry Index39.437.1
Released—2024-01-25
WeightsOpenProprietary
Context window—1M
Max output—33K
Input $ / M tokens—$0.40
Output $ / M tokens—$1.20
Results tracked1620

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Olmo 3.1 32b Instruct: 39.5 (#157), Qwen Plus: 38.9 (#167)

Coding benchmarks
BenchmarkOlmo 3.1 32b InstructQwen Plus
LMArena Coding13471328

Reasoning Qwen Plus leads

Olmo 3.1 32b Instruct: 26.4 (#132), Qwen Plus: 28.4 (#107)

Reasoning benchmarks
BenchmarkOlmo 3.1 32b InstructQwen Plus
LMArena Hard Prompts13221317
Kagi LLM Benchmark—63.3%
DTBench—81.1%
LMCA—24%

Math Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 36.3 (#167), Qwen Plus: 23.3 (#271)

Math benchmarks
BenchmarkOlmo 3.1 32b InstructQwen Plus
LMArena Math13051326
OTIS Mock AIME 2024-2025—17.8%
MATH Level 5—65.3%
FrontierMath (Feb 2025 set)—1.7%

Knowledge Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 36.1 (#175), Qwen Plus: 27.4 (#251)

Knowledge benchmarks
BenchmarkOlmo 3.1 32b InstructQwen Plus
LMArena Expert13081328
GPQA Diamond—48.1%

Multilingual Qwen Plus leads

Olmo 3.1 32b Instruct: 42.6 (#191), Qwen Plus: 45.1 (#175)

Multilingual benchmarks
BenchmarkOlmo 3.1 32b InstructQwen Plus
LMArena Non-English12751310
LMArena Chinese13041347
LMArena Russian12681323
LMArena French1328—
LMArena German1282—
LMArena Japanese—1251
LMArena Korean1206—
LMArena Spanish1336—

Instruction Following Too close to call

Olmo 3.1 32b Instruct: 68.6 (#187), Qwen Plus: 68.8 (#181)

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b InstructQwen Plus
LMArena Instruction Following12991303

Long Context Too close to call

Olmo 3.1 32b Instruct: 39.9 (#166), Qwen Plus: 40.3 (#158)

Long Context benchmarks
BenchmarkOlmo 3.1 32b InstructQwen Plus
LMArena Longer Query13121324

Writing & Preference Qwen Plus leads

Olmo 3.1 32b Instruct: 50.2 (#185), Qwen Plus: 52.2 (#176)

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b InstructQwen Plus
LMArena Text13111326
LMArena Creative Writing12641293
LMArena Multi-Turn13091336

Frequently asked questions

Is Olmo 3.1 32b Instruct better than Qwen Plus?

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 37.1 on the Noometry Index.

Is Olmo 3.1 32b Instruct or Qwen Plus better for coding?

They score almost the same on coding (39.5 vs 38.9); test both on your own repository before choosing.

How many benchmarks do Olmo 3.1 32b Instruct and Qwen Plus share?

12 benchmarks have published results for both models. Olmo 3.1 32b Instruct has 16 scored results on Noometry and Qwen Plus has 20.

Related comparisons

Go deeper