Model comparison

Olmo 3.1 32b Instruct vs Qwen3 8B

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 33.7 on the Noometry Index.

Last verified . 0 shared benchmarks.

Qwen3 8B Alibaba (Qwen)

33.7

Rank #238 Confirmed

Summary

  • The widest gap is in reasoning, where Olmo 3.1 32b Instruct leads 26.4 to 16.6.

Side by side

Olmo 3.1 32b Instruct and Qwen3 8B specifications
Olmo 3.1 32b InstructQwen3 8B
ProviderAllen Institute for AI (Ai2)Alibaba (Qwen)
Noometry Index39.433.7
Released—2025-04
WeightsOpenOpen
Context window—131K
Max output—8K
Input $ / M tokens—$0.18
Output $ / M tokens—$0.70
Results tracked1611

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 39.5 (#157), Qwen3 8B: 34.0 (#248)

Coding benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 8B
SciCode—22.6%
LMArena Coding1347—

Agentic & Tool Use Not comparable

Olmo 3.1 32b Instruct: —, Qwen3 8B: 30.2 (#78)

Agentic & Tool Use benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 8B
Berkeley Function Calling Leaderboard—42.6%

Reasoning Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 26.4 (#132), Qwen3 8B: 16.6 (#303)

Reasoning benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 8B
CritPt—0%
Chess Puzzles—5%
LMArena Hard Prompts1322—
DTBench—59.7%
LMCA—8.8%
Epoch Capabilities Index—136.17

Math Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 36.3 (#167), Qwen3 8B: 34.9 (#191)

Math benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 8B
OTIS Mock AIME 2024-2025—56.1%
LMArena Math1305—

Knowledge Too close to call

Olmo 3.1 32b Instruct: 36.1 (#175), Qwen3 8B: 36.1 (#173)

Knowledge benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 8B
GPQA Diamond—56.8%
Vectara Hallucination Rate—4.8%
LMArena Expert1308—

Multilingual Not comparable

Olmo 3.1 32b Instruct: 42.6 (#191), Qwen3 8B: —

Multilingual benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 8B
LMArena Non-English1275—
LMArena Chinese1304—
LMArena French1328—
LMArena German1282—
LMArena Korean1206—
LMArena Russian1268—
LMArena Spanish1336—

Instruction Following Not comparable

Olmo 3.1 32b Instruct: 68.6 (#187), Qwen3 8B: —

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 8B
LMArena Instruction Following1299—

Long Context Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 39.9 (#166), Qwen3 8B: 37.9 (#210)

Long Context benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 8B
Fiction.LiveBench—62.1%
LMArena Longer Query1312—

Writing & Preference Not comparable

Olmo 3.1 32b Instruct: 50.2 (#185), Qwen3 8B: —

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 8B
LMArena Text1311—
LMArena Creative Writing1264—
LMArena Multi-Turn1309—

Frequently asked questions

Is Olmo 3.1 32b Instruct better than Qwen3 8B?

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 33.7 on the Noometry Index.

Is Olmo 3.1 32b Instruct or Qwen3 8B better for coding?

Olmo 3.1 32b Instruct scores higher on coding benchmarks: 39.5 versus 34.0 in the Noometry coding category.

How many benchmarks do Olmo 3.1 32b Instruct and Qwen3 8B share?

0 benchmarks have published results for both models. Olmo 3.1 32b Instruct has 16 scored results on Noometry and Qwen3 8B has 11.

Related comparisons

Go deeper