Model comparison

Olmo 3.1 32b Instruct vs Qwen3 Coder Next

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 34.3 on the Noometry Index.

Last verified . 0 shared benchmarks.

Summary

  • The widest gap is in reasoning, where Olmo 3.1 32b Instruct leads 26.4 to 22.4.

Side by side

Olmo 3.1 32b Instruct and Qwen3 Coder Next specifications
Olmo 3.1 32b InstructQwen3 Coder Next
ProviderAllen Institute for AI (Ai2)Alibaba (Qwen)
Noometry Index39.434.3
Released—2026-02-02
WeightsOpenOpen
Context window—262K
Max output—66K
Input $ / M tokens—$0.12
Output $ / M tokens—$0.80
Results tracked163

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 39.5 (#157), Qwen3 Coder Next: 36.3 (#210)

Coding benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 Coder Next
SciCode—32.3%
WeirdML—34.4%
LMArena Coding1347—

Reasoning Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 26.4 (#132), Qwen3 Coder Next: 22.4 (#196)

Reasoning benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 Coder Next
CritPt—0%
LMArena Hard Prompts1322—

Math Not comparable

Olmo 3.1 32b Instruct: 36.3 (#167), Qwen3 Coder Next: —

Math benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 Coder Next
LMArena Math1305—

Knowledge Not comparable

Olmo 3.1 32b Instruct: 36.1 (#175), Qwen3 Coder Next: —

Knowledge benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 Coder Next
LMArena Expert1308—

Multilingual Not comparable

Olmo 3.1 32b Instruct: 42.6 (#191), Qwen3 Coder Next: —

Multilingual benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 Coder Next
LMArena Non-English1275—
LMArena Chinese1304—
LMArena French1328—
LMArena German1282—
LMArena Korean1206—
LMArena Russian1268—
LMArena Spanish1336—

Instruction Following Not comparable

Olmo 3.1 32b Instruct: 68.6 (#187), Qwen3 Coder Next: —

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 Coder Next
LMArena Instruction Following1299—

Long Context Not comparable

Olmo 3.1 32b Instruct: 39.9 (#166), Qwen3 Coder Next: —

Long Context benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 Coder Next
LMArena Longer Query1312—

Writing & Preference Not comparable

Olmo 3.1 32b Instruct: 50.2 (#185), Qwen3 Coder Next: —

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 Coder Next
LMArena Text1311—
LMArena Creative Writing1264—
LMArena Multi-Turn1309—

Frequently asked questions

Is Olmo 3.1 32b Instruct better than Qwen3 Coder Next?

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 34.3 on the Noometry Index.

Is Olmo 3.1 32b Instruct or Qwen3 Coder Next better for coding?

Olmo 3.1 32b Instruct scores higher on coding benchmarks: 39.5 versus 36.3 in the Noometry coding category.

How many benchmarks do Olmo 3.1 32b Instruct and Qwen3 Coder Next share?

0 benchmarks have published results for both models. Olmo 3.1 32b Instruct has 16 scored results on Noometry and Qwen3 Coder Next has 3.

Related comparisons

Go deeper