Model comparison

Olmo 3.1 32b Think vs Qwen2.5 Plus 1127

Olmo 3.1 32b Think and Qwen2.5 Plus 1127 score almost the same on the Noometry Index (37.9 vs 38.8), so choose on price, context window or the category you care about most.

Last verified . 13 shared benchmarks.

Summary

  • They share 13 benchmarks with published results for both. Olmo 3.1 32b Think scores higher in 2 categories and Qwen2.5 Plus 1127 in 6 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Qwen2.5 Plus 1127 leads 41.9 to 38.1.
  • Olmo 3.1 32b Think has downloadable open weights; the other is API-only.

Side by side

Olmo 3.1 32b Think and Qwen2.5 Plus 1127 specifications
Olmo 3.1 32b ThinkQwen2.5 Plus 1127
ProviderAllen Institute for AI (Ai2)Alibaba (Qwen)
Noometry Index37.938.8
Released——
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1514

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Olmo 3.1 32b Think: 37.7 (#189), Qwen2.5 Plus 1127: 38.5 (#175)

Coding benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen2.5 Plus 1127
LMArena Coding12911314

Reasoning Too close to call

Olmo 3.1 32b Think: 25.2 (#150), Qwen2.5 Plus 1127: 25.9 (#141)

Reasoning benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen2.5 Plus 1127
LMArena Hard Prompts12721299

Math Too close to call

Olmo 3.1 32b Think: 36.3 (#168), Qwen2.5 Plus 1127: 36.1 (#174)

Math benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen2.5 Plus 1127
LMArena Math13051298

Knowledge Too close to call

Olmo 3.1 32b Think: 35.7 (#181), Qwen2.5 Plus 1127: 35.5 (#183)

Knowledge benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen2.5 Plus 1127
LMArena Expert12951289

Multilingual Qwen2.5 Plus 1127 leads

Olmo 3.1 32b Think: 38.1 (#231), Qwen2.5 Plus 1127: 41.9 (#201)

Multilingual benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen2.5 Plus 1127
LMArena Non-English12091265
LMArena Chinese12421314
LMArena German12621231
LMArena Russian11931271
LMArena French1260—
LMArena Japanese—1207
LMArena Spanish1289—

Instruction Following Qwen2.5 Plus 1127 leads

Olmo 3.1 32b Think: 65.6 (#218), Qwen2.5 Plus 1127: 67.2 (#199)

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen2.5 Plus 1127
LMArena Instruction Following12471275

Long Context Too close to call

Olmo 3.1 32b Think: 38.6 (#195), Qwen2.5 Plus 1127: 39.2 (#184)

Long Context benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen2.5 Plus 1127
LMArena Longer Query12721292

Writing & Preference Qwen2.5 Plus 1127 leads

Olmo 3.1 32b Think: 46.2 (#220), Qwen2.5 Plus 1127: 49.4 (#192)

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen2.5 Plus 1127
LMArena Text12721299
LMArena Creative Writing12261262
LMArena Multi-Turn12521299

Frequently asked questions

Is Olmo 3.1 32b Think better than Qwen2.5 Plus 1127?

Olmo 3.1 32b Think and Qwen2.5 Plus 1127 score almost the same on the Noometry Index (37.9 vs 38.8), so choose on price, context window or the category you care about most.

Is Olmo 3.1 32b Think or Qwen2.5 Plus 1127 better for coding?

They score almost the same on coding (37.7 vs 38.5); test both on your own repository before choosing.

How many benchmarks do Olmo 3.1 32b Think and Qwen2.5 Plus 1127 share?

13 benchmarks have published results for both models. Olmo 3.1 32b Think has 15 scored results on Noometry and Qwen2.5 Plus 1127 has 14.

Related comparisons

Go deeper