Model comparison

Olmo 3.1 32b Instruct vs Qwen2.5-Coder (1.5B)

Olmo 3.1 32b Instruct has enough public results to be ranked (#168); Qwen2.5-Coder (1.5B) does not yet, so treat this comparison as directional.

Last verified . 0 shared benchmarks.

Side by side

Olmo 3.1 32b Instruct and Qwen2.5-Coder (1.5B) specifications
Olmo 3.1 32b InstructQwen2.5-Coder (1.5B)
ProviderAllen Institute for AI (Ai2)Alibaba (Qwen)
Noometry Index39.4—
Released—2024-09-18
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked166

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Olmo 3.1 32b Instruct: 39.5 (#157), Qwen2.5-Coder (1.5B): —

Coding benchmarks
BenchmarkOlmo 3.1 32b InstructQwen2.5-Coder (1.5B)
LMArena Coding1347—

Reasoning Not comparable

Olmo 3.1 32b Instruct: 26.4 (#132), Qwen2.5-Coder (1.5B): —

Reasoning benchmarks
BenchmarkOlmo 3.1 32b InstructQwen2.5-Coder (1.5B)
LMArena Hard Prompts1322—
Epoch Capabilities Index—113.14
HellaSwag—76.8%
WinoGrande—72.9%

Math Not comparable

Olmo 3.1 32b Instruct: 36.3 (#167), Qwen2.5-Coder (1.5B): —

Math benchmarks
BenchmarkOlmo 3.1 32b InstructQwen2.5-Coder (1.5B)
LMArena Math1305—
GSM8K—86.7%

Knowledge Not comparable

Olmo 3.1 32b Instruct: 36.1 (#175), Qwen2.5-Coder (1.5B): —

Knowledge benchmarks
BenchmarkOlmo 3.1 32b InstructQwen2.5-Coder (1.5B)
LMArena Expert1308—
ARC (AI2) Challenge—60.9%
MMLU—68%

Multilingual Not comparable

Olmo 3.1 32b Instruct: 42.6 (#191), Qwen2.5-Coder (1.5B): —

Multilingual benchmarks
BenchmarkOlmo 3.1 32b InstructQwen2.5-Coder (1.5B)
LMArena Non-English1275—
LMArena Chinese1304—
LMArena French1328—
LMArena German1282—
LMArena Korean1206—
LMArena Russian1268—
LMArena Spanish1336—

Instruction Following Not comparable

Olmo 3.1 32b Instruct: 68.6 (#187), Qwen2.5-Coder (1.5B): —

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b InstructQwen2.5-Coder (1.5B)
LMArena Instruction Following1299—

Long Context Not comparable

Olmo 3.1 32b Instruct: 39.9 (#166), Qwen2.5-Coder (1.5B): —

Long Context benchmarks
BenchmarkOlmo 3.1 32b InstructQwen2.5-Coder (1.5B)
LMArena Longer Query1312—

Writing & Preference Not comparable

Olmo 3.1 32b Instruct: 50.2 (#185), Qwen2.5-Coder (1.5B): —

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b InstructQwen2.5-Coder (1.5B)
LMArena Text1311—
LMArena Creative Writing1264—
LMArena Multi-Turn1309—

Frequently asked questions

Is Olmo 3.1 32b Instruct better than Qwen2.5-Coder (1.5B)?

Olmo 3.1 32b Instruct has enough public results to be ranked (#168); Qwen2.5-Coder (1.5B) does not yet, so treat this comparison as directional.

How many benchmarks do Olmo 3.1 32b Instruct and Qwen2.5-Coder (1.5B) share?

0 benchmarks have published results for both models. Olmo 3.1 32b Instruct has 16 scored results on Noometry and Qwen2.5-Coder (1.5B) has 6.

Related comparisons

Go deeper