Model comparison

Olmo 3.1 32b Instruct vs Qwen3.7 Plus

Qwen3.7 Plus is the stronger model overall, scoring 45.3 to 39.4 on the Noometry Index.

Last verified . 16 shared benchmarks.

Summary

  • They share 16 benchmarks with published results for both. Olmo 3.1 32b Instruct scores higher in 1 category and Qwen3.7 Plus in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3.7 Plus leads 54.9 to 36.1.
  • Olmo 3.1 32b Instruct has downloadable open weights; the other is API-only.

Side by side

Olmo 3.1 32b Instruct and Qwen3.7 Plus specifications
Olmo 3.1 32b InstructQwen3.7 Plus
ProviderAllen Institute for AI (Ai2)Alibaba (Qwen)
Noometry Index39.445.3
Released—2026-06-02
WeightsOpenProprietary
Context window—1M
Max output—131K
Input $ / M tokens—$0.40
Output $ / M tokens—$1.60
Results tracked1632

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 39.5 (#157), Qwen3.7 Plus: 36.6 (#206)

Coding benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3.7 Plus
LMArena Coding13471473
FrontierCode—10.2%
SciCode—45.5%

Agentic & Tool Use Not comparable

Olmo 3.1 32b Instruct: —, Qwen3.7 Plus: 21.4 (#138)

Agentic & Tool Use benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3.7 Plus
OSWorld 2.0—2.8%

Reasoning Qwen3.7 Plus leads

Olmo 3.1 32b Instruct: 26.4 (#132), Qwen3.7 Plus: 39.3 (#59)

Reasoning benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3.7 Plus
LMArena Hard Prompts13221460
NYT Connections (extended)—74.8%
CritPt—9.1%
Chess Puzzles—24%
Mystery Game Puzzles—17%
DTBench—84%
LMCA—37.6%
Epoch Capabilities Index—147.37

Math Qwen3.7 Plus leads

Olmo 3.1 32b Instruct: 36.3 (#167), Qwen3.7 Plus: 50.5 (#56)

Math benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3.7 Plus
LMArena Math13051466
FrontierMath (Tiers 1-3)—34.4%
OTIS Mock AIME 2024-2025—93.3%

Knowledge Qwen3.7 Plus leads

Olmo 3.1 32b Instruct: 36.1 (#175), Qwen3.7 Plus: 54.9 (#51)

Knowledge benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3.7 Plus
LMArena Expert13081467
GPQA Diamond—87.9%

Multimodal Not comparable

Olmo 3.1 32b Instruct: —, Qwen3.7 Plus: 41.8 (#33)

Multimodal benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3.7 Plus
LMArena Vision—1279
LMArena Document—1444

Multilingual Qwen3.7 Plus leads

Olmo 3.1 32b Instruct: 42.6 (#191), Qwen3.7 Plus: 54.8 (#38)

Multilingual benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3.7 Plus
LMArena Non-English12751445
LMArena Chinese13041510
LMArena French13281473
LMArena German12821471
LMArena Korean12061415
LMArena Russian12681457
LMArena Spanish13361457
LMArena Japanese—1413

Instruction Following Qwen3.7 Plus leads

Olmo 3.1 32b Instruct: 68.6 (#187), Qwen3.7 Plus: 75.8 (#52)

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3.7 Plus
LMArena Instruction Following12991440

Long Context Qwen3.7 Plus leads

Olmo 3.1 32b Instruct: 39.9 (#166), Qwen3.7 Plus: 44.5 (#65)

Long Context benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3.7 Plus
LMArena Longer Query13121455

Writing & Preference Qwen3.7 Plus leads

Olmo 3.1 32b Instruct: 50.2 (#185), Qwen3.7 Plus: 64.3 (#56)

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3.7 Plus
LMArena Text13111455
LMArena Creative Writing12641439
LMArena Multi-Turn13091460

Frequently asked questions

Is Olmo 3.1 32b Instruct better than Qwen3.7 Plus?

Qwen3.7 Plus is the stronger model overall, scoring 45.3 to 39.4 on the Noometry Index.

Is Olmo 3.1 32b Instruct or Qwen3.7 Plus better for coding?

Olmo 3.1 32b Instruct scores higher on coding benchmarks: 39.5 versus 36.6 in the Noometry coding category.

How many benchmarks do Olmo 3.1 32b Instruct and Qwen3.7 Plus share?

16 benchmarks have published results for both models. Olmo 3.1 32b Instruct has 16 scored results on Noometry and Qwen3.7 Plus has 32.

Related comparisons

Go deeper