Model comparison

Olmo 3.1 32b Instruct vs Qwen2.5 Plus 1127

Olmo 3.1 32b Instruct and Qwen2.5 Plus 1127 score almost the same on the Noometry Index (39.4 vs 38.8), so choose on price, context window or the category you care about most.

Last verified . 13 shared benchmarks.

Summary

  • They share 13 benchmarks with published results for both. Olmo 3.1 32b Instruct scores higher in 8 categories and Qwen2.5 Plus 1127 in 0 categories; 2 gaps are clear of the uncertainty.
  • Olmo 3.1 32b Instruct has downloadable open weights; the other is API-only.

Side by side

Olmo 3.1 32b Instruct and Qwen2.5 Plus 1127 specifications
Olmo 3.1 32b InstructQwen2.5 Plus 1127
ProviderAllen Institute for AI (Ai2)Alibaba (Qwen)
Noometry Index39.438.8
Released——
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1614

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 39.5 (#157), Qwen2.5 Plus 1127: 38.5 (#175)

Coding benchmarks
BenchmarkOlmo 3.1 32b InstructQwen2.5 Plus 1127
LMArena Coding13471314

Reasoning Too close to call

Olmo 3.1 32b Instruct: 26.4 (#132), Qwen2.5 Plus 1127: 25.9 (#141)

Reasoning benchmarks
BenchmarkOlmo 3.1 32b InstructQwen2.5 Plus 1127
LMArena Hard Prompts13221299

Math Too close to call

Olmo 3.1 32b Instruct: 36.3 (#167), Qwen2.5 Plus 1127: 36.1 (#174)

Math benchmarks
BenchmarkOlmo 3.1 32b InstructQwen2.5 Plus 1127
LMArena Math13051298

Knowledge Too close to call

Olmo 3.1 32b Instruct: 36.1 (#175), Qwen2.5 Plus 1127: 35.5 (#183)

Knowledge benchmarks
BenchmarkOlmo 3.1 32b InstructQwen2.5 Plus 1127
LMArena Expert13081289

Multilingual Too close to call

Olmo 3.1 32b Instruct: 42.6 (#191), Qwen2.5 Plus 1127: 41.9 (#201)

Multilingual benchmarks
BenchmarkOlmo 3.1 32b InstructQwen2.5 Plus 1127
LMArena Non-English12751265
LMArena Chinese13041314
LMArena German12821231
LMArena Russian12681271
LMArena French1328—
LMArena Japanese—1207
LMArena Korean1206—
LMArena Spanish1336—

Instruction Following Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 68.6 (#187), Qwen2.5 Plus 1127: 67.2 (#199)

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b InstructQwen2.5 Plus 1127
LMArena Instruction Following12991275

Long Context Too close to call

Olmo 3.1 32b Instruct: 39.9 (#166), Qwen2.5 Plus 1127: 39.2 (#184)

Long Context benchmarks
BenchmarkOlmo 3.1 32b InstructQwen2.5 Plus 1127
LMArena Longer Query13121292

Writing & Preference Too close to call

Olmo 3.1 32b Instruct: 50.2 (#185), Qwen2.5 Plus 1127: 49.4 (#192)

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b InstructQwen2.5 Plus 1127
LMArena Text13111299
LMArena Creative Writing12641262
LMArena Multi-Turn13091299

Frequently asked questions

Is Olmo 3.1 32b Instruct better than Qwen2.5 Plus 1127?

Olmo 3.1 32b Instruct and Qwen2.5 Plus 1127 score almost the same on the Noometry Index (39.4 vs 38.8), so choose on price, context window or the category you care about most.

Is Olmo 3.1 32b Instruct or Qwen2.5 Plus 1127 better for coding?

Olmo 3.1 32b Instruct scores higher on coding benchmarks: 39.5 versus 38.5 in the Noometry coding category.

How many benchmarks do Olmo 3.1 32b Instruct and Qwen2.5 Plus 1127 share?

13 benchmarks have published results for both models. Olmo 3.1 32b Instruct has 16 scored results on Noometry and Qwen2.5 Plus 1127 has 14.

Related comparisons

Go deeper