Model comparison

Olmo 3 32b Think vs Qwen2.5 Plus 1127

Olmo 3 32b Think and Qwen2.5 Plus 1127 score almost the same on the Noometry Index (38.7 vs 38.8), so choose on price, context window or the category you care about most.

Last verified . 13 shared benchmarks.

Summary

  • They share 13 benchmarks with published results for both. Olmo 3 32b Think scores higher in 5 categories and Qwen2.5 Plus 1127 in 3 categories, but none of those gaps is larger than the uncertainty.
  • Olmo 3 32b Think has downloadable open weights; the other is API-only.

Side by side

Olmo 3 32b Think and Qwen2.5 Plus 1127 specifications
Olmo 3 32b ThinkQwen2.5 Plus 1127
ProviderAllen Institute for AI (Ai2)Alibaba (Qwen)
Noometry Index38.738.8
Released——
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1414

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Olmo 3 32b Think: 38.6 (#172), Qwen2.5 Plus 1127: 38.5 (#175)

Coding benchmarks
BenchmarkOlmo 3 32b ThinkQwen2.5 Plus 1127
LMArena Coding13191314

Reasoning Too close to call

Olmo 3 32b Think: 25.9 (#140), Qwen2.5 Plus 1127: 25.9 (#141)

Reasoning benchmarks
BenchmarkOlmo 3 32b ThinkQwen2.5 Plus 1127
LMArena Hard Prompts13021299

Math Too close to call

Olmo 3 32b Think: 36.5 (#165), Qwen2.5 Plus 1127: 36.1 (#174)

Math benchmarks
BenchmarkOlmo 3 32b ThinkQwen2.5 Plus 1127
LMArena Math13161298

Knowledge Too close to call

Olmo 3 32b Think: 35.0 (#190), Qwen2.5 Plus 1127: 35.5 (#183)

Knowledge benchmarks
BenchmarkOlmo 3 32b ThinkQwen2.5 Plus 1127
LMArena Expert12731289

Multilingual Too close to call

Olmo 3 32b Think: 41.2 (#210), Qwen2.5 Plus 1127: 41.9 (#201)

Multilingual benchmarks
BenchmarkOlmo 3 32b ThinkQwen2.5 Plus 1127
LMArena Non-English12551265
LMArena Chinese13001314
LMArena German12901231
LMArena Russian12541271
LMArena French1291—
LMArena Japanese—1207

Instruction Following Too close to call

Olmo 3 32b Think: 67.2 (#198), Qwen2.5 Plus 1127: 67.2 (#199)

Instruction Following benchmarks
BenchmarkOlmo 3 32b ThinkQwen2.5 Plus 1127
LMArena Instruction Following12751275

Long Context Too close to call

Olmo 3 32b Think: 39.4 (#182), Qwen2.5 Plus 1127: 39.2 (#184)

Long Context benchmarks
BenchmarkOlmo 3 32b ThinkQwen2.5 Plus 1127
LMArena Longer Query12961292

Writing & Preference Too close to call

Olmo 3 32b Think: 49.1 (#193), Qwen2.5 Plus 1127: 49.4 (#192)

Writing & Preference benchmarks
BenchmarkOlmo 3 32b ThinkQwen2.5 Plus 1127
LMArena Text13001299
LMArena Creative Writing12561262
LMArena Multi-Turn12901299

Frequently asked questions

Is Olmo 3 32b Think better than Qwen2.5 Plus 1127?

Olmo 3 32b Think and Qwen2.5 Plus 1127 score almost the same on the Noometry Index (38.7 vs 38.8), so choose on price, context window or the category you care about most.

Is Olmo 3 32b Think or Qwen2.5 Plus 1127 better for coding?

They score almost the same on coding (38.6 vs 38.5); test both on your own repository before choosing.

How many benchmarks do Olmo 3 32b Think and Qwen2.5 Plus 1127 share?

13 benchmarks have published results for both models. Olmo 3 32b Think has 14 scored results on Noometry and Qwen2.5 Plus 1127 has 14.

Related comparisons

Go deeper