Model comparison

Olmo 7b Instruct vs Qwen1.5 4b Chat

Olmo 7b Instruct is the stronger model overall, scoring 30.3 to 28.8 on the Noometry Index.

Last verified . 10 shared benchmarks.

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Olmo 7b Instruct scores higher in 3 categories and Qwen1.5 4b Chat in 3 categories; one gap is clear of the uncertainty.

Side by side

Olmo 7b Instruct and Qwen1.5 4b Chat specifications
Olmo 7b InstructQwen1.5 4b Chat
ProviderAllen Institute for AI (Ai2)Alibaba (Qwen)
Noometry Index30.328.8
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1013

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Olmo 7b Instruct: 29.6 (#303), Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkOlmo 7b InstructQwen1.5 4b Chat
LMArena Coding1016999

Reasoning Too close to call

Olmo 7b Instruct: 18.8 (#274), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkOlmo 7b InstructQwen1.5 4b Chat
LMArena Hard Prompts993976

Math Too close to call

Olmo 7b Instruct: 30.2 (#237), Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkOlmo 7b InstructQwen1.5 4b Chat
LMArena Math10181026

Knowledge Not comparable

Olmo 7b Instruct: —, Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkOlmo 7b InstructQwen1.5 4b Chat
LMArena Expert—980

Multilingual Too close to call

Olmo 7b Instruct: 24.0 (#291), Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkOlmo 7b InstructQwen1.5 4b Chat
LMArena Non-English977979
LMArena Chinese10141024
LMArena Russian947952
LMArena German—902

Instruction Following Too close to call

Olmo 7b Instruct: 49.0 (#301), Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkOlmo 7b InstructQwen1.5 4b Chat
LMArena Instruction Following978978

Long Context Not comparable

Olmo 7b Instruct: —, Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkOlmo 7b InstructQwen1.5 4b Chat
LMArena Longer Query—988

Writing & Preference Olmo 7b Instruct leads

Olmo 7b Instruct: 25.8 (#303), Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkOlmo 7b InstructQwen1.5 4b Chat
LMArena Text1032997
LMArena Creative Writing990969
LMArena Multi-Turn1007977

Frequently asked questions

Is Olmo 7b Instruct better than Qwen1.5 4b Chat?

Olmo 7b Instruct is the stronger model overall, scoring 30.3 to 28.8 on the Noometry Index.

Is Olmo 7b Instruct or Qwen1.5 4b Chat better for coding?

They score almost the same on coding (29.6 vs 29.1); test both on your own repository before choosing.

How many benchmarks do Olmo 7b Instruct and Qwen1.5 4b Chat share?

10 benchmarks have published results for both models. Olmo 7b Instruct has 10 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper