Model comparison
Olmo 3.1 32b Instruct vs Qwen3 8B
Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 33.7 on the Noometry Index.
Last verified . 0 shared benchmarks.
Summary
- The widest gap is in reasoning, where Olmo 3.1 32b Instruct leads 26.4 to 16.6.
Side by side
| Olmo 3.1 32b Instruct | Qwen3 8B | |
|---|---|---|
| Provider | Allen Institute for AI (Ai2) | Alibaba (Qwen) |
| Noometry Index | 39.4 | 33.7 |
| Released | — | 2025-04 |
| Weights | Open | Open |
| Context window | — | 131K |
| Max output | — | 8K |
| Input $ / M tokens | — | $0.18 |
| Output $ / M tokens | — | $0.70 |
| Results tracked | 16 | 11 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Olmo 3.1 32b Instruct leads
Olmo 3.1 32b Instruct: 39.5 (#157), Qwen3 8B: 34.0 (#248)
| Benchmark | Olmo 3.1 32b Instruct | Qwen3 8B |
|---|---|---|
| SciCode | — | 22.6% |
| LMArena Coding | 1347 | — |
Agentic & Tool Use Not comparable
Olmo 3.1 32b Instruct: —, Qwen3 8B: 30.2 (#78)
| Benchmark | Olmo 3.1 32b Instruct | Qwen3 8B |
|---|---|---|
| Berkeley Function Calling Leaderboard | — | 42.6% |
Reasoning Olmo 3.1 32b Instruct leads
Olmo 3.1 32b Instruct: 26.4 (#132), Qwen3 8B: 16.6 (#303)
| Benchmark | Olmo 3.1 32b Instruct | Qwen3 8B |
|---|---|---|
| CritPt | — | 0% |
| Chess Puzzles | — | 5% |
| LMArena Hard Prompts | 1322 | — |
| DTBench | — | 59.7% |
| LMCA | — | 8.8% |
| Epoch Capabilities Index | — | 136.17 |
Math Olmo 3.1 32b Instruct leads
Olmo 3.1 32b Instruct: 36.3 (#167), Qwen3 8B: 34.9 (#191)
| Benchmark | Olmo 3.1 32b Instruct | Qwen3 8B |
|---|---|---|
| OTIS Mock AIME 2024-2025 | — | 56.1% |
| LMArena Math | 1305 | — |
Knowledge Too close to call
Olmo 3.1 32b Instruct: 36.1 (#175), Qwen3 8B: 36.1 (#173)
| Benchmark | Olmo 3.1 32b Instruct | Qwen3 8B |
|---|---|---|
| GPQA Diamond | — | 56.8% |
| Vectara Hallucination Rate | — | 4.8% |
| LMArena Expert | 1308 | — |
Multilingual Not comparable
Olmo 3.1 32b Instruct: 42.6 (#191), Qwen3 8B: —
| Benchmark | Olmo 3.1 32b Instruct | Qwen3 8B |
|---|---|---|
| LMArena Non-English | 1275 | — |
| LMArena Chinese | 1304 | — |
| LMArena French | 1328 | — |
| LMArena German | 1282 | — |
| LMArena Korean | 1206 | — |
| LMArena Russian | 1268 | — |
| LMArena Spanish | 1336 | — |
Instruction Following Not comparable
Olmo 3.1 32b Instruct: 68.6 (#187), Qwen3 8B: —
| Benchmark | Olmo 3.1 32b Instruct | Qwen3 8B |
|---|---|---|
| LMArena Instruction Following | 1299 | — |
Long Context Olmo 3.1 32b Instruct leads
Olmo 3.1 32b Instruct: 39.9 (#166), Qwen3 8B: 37.9 (#210)
| Benchmark | Olmo 3.1 32b Instruct | Qwen3 8B |
|---|---|---|
| Fiction.LiveBench | — | 62.1% |
| LMArena Longer Query | 1312 | — |
Writing & Preference Not comparable
Olmo 3.1 32b Instruct: 50.2 (#185), Qwen3 8B: —
| Benchmark | Olmo 3.1 32b Instruct | Qwen3 8B |
|---|---|---|
| LMArena Text | 1311 | — |
| LMArena Creative Writing | 1264 | — |
| LMArena Multi-Turn | 1309 | — |
Frequently asked questions
Is Olmo 3.1 32b Instruct better than Qwen3 8B?
Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 33.7 on the Noometry Index.
Is Olmo 3.1 32b Instruct or Qwen3 8B better for coding?
Olmo 3.1 32b Instruct scores higher on coding benchmarks: 39.5 versus 34.0 in the Noometry coding category.
How many benchmarks do Olmo 3.1 32b Instruct and Qwen3 8B share?
0 benchmarks have published results for both models. Olmo 3.1 32b Instruct has 16 scored results on Noometry and Qwen3 8B has 11.