Model comparison

Olmo 3.1 32b Think vs Qwen3 Coder Next

Olmo 3.1 32b Think is the stronger model overall, scoring 37.9 to 34.3 on the Noometry Index.

Last verified . 0 shared benchmarks.

Side by side

Olmo 3.1 32b Think and Qwen3 Coder Next specifications
Olmo 3.1 32b ThinkQwen3 Coder Next
ProviderAllen Institute for AI (Ai2)Alibaba (Qwen)
Noometry Index37.934.3
Released—2026-02-02
WeightsOpenOpen
Context window—262K
Max output—66K
Input $ / M tokens—$0.12
Output $ / M tokens—$0.80
Results tracked153

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Think leads

Olmo 3.1 32b Think: 37.7 (#189), Qwen3 Coder Next: 36.3 (#210)

Coding benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen3 Coder Next
SciCode—32.3%
WeirdML—34.4%
LMArena Coding1291—

Reasoning Olmo 3.1 32b Think leads

Olmo 3.1 32b Think: 25.2 (#150), Qwen3 Coder Next: 22.4 (#196)

Reasoning benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen3 Coder Next
CritPt—0%
LMArena Hard Prompts1272—

Math Not comparable

Olmo 3.1 32b Think: 36.3 (#168), Qwen3 Coder Next: —

Math benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen3 Coder Next
LMArena Math1305—

Knowledge Not comparable

Olmo 3.1 32b Think: 35.7 (#181), Qwen3 Coder Next: —

Knowledge benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen3 Coder Next
LMArena Expert1295—

Multilingual Not comparable

Olmo 3.1 32b Think: 38.1 (#231), Qwen3 Coder Next: —

Multilingual benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen3 Coder Next
LMArena Non-English1209—
LMArena Chinese1242—
LMArena French1260—
LMArena German1262—
LMArena Russian1193—
LMArena Spanish1289—

Instruction Following Not comparable

Olmo 3.1 32b Think: 65.6 (#218), Qwen3 Coder Next: —

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen3 Coder Next
LMArena Instruction Following1247—

Long Context Not comparable

Olmo 3.1 32b Think: 38.6 (#195), Qwen3 Coder Next: —

Long Context benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen3 Coder Next
LMArena Longer Query1272—

Writing & Preference Not comparable

Olmo 3.1 32b Think: 46.2 (#220), Qwen3 Coder Next: —

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen3 Coder Next
LMArena Text1272—
LMArena Creative Writing1226—
LMArena Multi-Turn1252—

Frequently asked questions

Is Olmo 3.1 32b Think better than Qwen3 Coder Next?

Olmo 3.1 32b Think is the stronger model overall, scoring 37.9 to 34.3 on the Noometry Index.

Is Olmo 3.1 32b Think or Qwen3 Coder Next better for coding?

Olmo 3.1 32b Think scores higher on coding benchmarks: 37.7 versus 36.3 in the Noometry coding category.

How many benchmarks do Olmo 3.1 32b Think and Qwen3 Coder Next share?

0 benchmarks have published results for both models. Olmo 3.1 32b Think has 15 scored results on Noometry and Qwen3 Coder Next has 3.

Related comparisons

Go deeper