Model comparison

Olmo 7b Instruct vs Qwen2.5 Plus 1127

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 30.3 on the Noometry Index.

Last verified . 10 shared benchmarks.

Summary

  • They share 10 benchmarks with published results for both. Olmo 7b Instruct scores higher in 0 categories and Qwen2.5 Plus 1127 in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen2.5 Plus 1127 leads 49.4 to 25.8.
  • Olmo 7b Instruct has downloadable open weights; the other is API-only.

Side by side

Olmo 7b Instruct and Qwen2.5 Plus 1127 specifications
Olmo 7b InstructQwen2.5 Plus 1127
ProviderAllen Institute for AI (Ai2)Alibaba (Qwen)
Noometry Index30.338.8
Released——
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1014

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 Plus 1127 leads

Olmo 7b Instruct: 29.6 (#303), Qwen2.5 Plus 1127: 38.5 (#175)

Coding benchmarks
BenchmarkOlmo 7b InstructQwen2.5 Plus 1127
LMArena Coding10161314

Reasoning Qwen2.5 Plus 1127 leads

Olmo 7b Instruct: 18.8 (#274), Qwen2.5 Plus 1127: 25.9 (#141)

Reasoning benchmarks
BenchmarkOlmo 7b InstructQwen2.5 Plus 1127
LMArena Hard Prompts9931299

Math Qwen2.5 Plus 1127 leads

Olmo 7b Instruct: 30.2 (#237), Qwen2.5 Plus 1127: 36.1 (#174)

Math benchmarks
BenchmarkOlmo 7b InstructQwen2.5 Plus 1127
LMArena Math10181298

Knowledge Not comparable

Olmo 7b Instruct: —, Qwen2.5 Plus 1127: 35.5 (#183)

Knowledge benchmarks
BenchmarkOlmo 7b InstructQwen2.5 Plus 1127
LMArena Expert—1289

Multilingual Qwen2.5 Plus 1127 leads

Olmo 7b Instruct: 24.0 (#291), Qwen2.5 Plus 1127: 41.9 (#201)

Multilingual benchmarks
BenchmarkOlmo 7b InstructQwen2.5 Plus 1127
LMArena Non-English9771265
LMArena Chinese10141314
LMArena Russian9471271
LMArena German—1231
LMArena Japanese—1207

Instruction Following Qwen2.5 Plus 1127 leads

Olmo 7b Instruct: 49.0 (#301), Qwen2.5 Plus 1127: 67.2 (#199)

Instruction Following benchmarks
BenchmarkOlmo 7b InstructQwen2.5 Plus 1127
LMArena Instruction Following9781275

Long Context Not comparable

Olmo 7b Instruct: —, Qwen2.5 Plus 1127: 39.2 (#184)

Long Context benchmarks
BenchmarkOlmo 7b InstructQwen2.5 Plus 1127
LMArena Longer Query—1292

Writing & Preference Qwen2.5 Plus 1127 leads

Olmo 7b Instruct: 25.8 (#303), Qwen2.5 Plus 1127: 49.4 (#192)

Writing & Preference benchmarks
BenchmarkOlmo 7b InstructQwen2.5 Plus 1127
LMArena Text10321299
LMArena Creative Writing9901262
LMArena Multi-Turn10071299

Frequently asked questions

Is Olmo 7b Instruct better than Qwen2.5 Plus 1127?

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 30.3 on the Noometry Index.

Is Olmo 7b Instruct or Qwen2.5 Plus 1127 better for coding?

Qwen2.5 Plus 1127 scores higher on coding benchmarks: 38.5 versus 29.6 in the Noometry coding category.

How many benchmarks do Olmo 7b Instruct and Qwen2.5 Plus 1127 share?

10 benchmarks have published results for both models. Olmo 7b Instruct has 10 scored results on Noometry and Qwen2.5 Plus 1127 has 14.

Related comparisons

Go deeper