Model comparison

Olmo 3.1 32b Instruct vs Qwen2.5 72B Instruct

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 31.9 on the Noometry Index.

Last verified . 16 shared benchmarks.

Summary

  • They share 16 benchmarks with published results for both. Olmo 3.1 32b Instruct scores higher in 8 categories and Qwen2.5 72B Instruct in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Olmo 3.1 32b Instruct leads 36.3 to 19.3.

Side by side

Olmo 3.1 32b Instruct and Qwen2.5 72B Instruct specifications
Olmo 3.1 32b InstructQwen2.5 72B Instruct
ProviderAllen Institute for AI (Ai2)Alibaba (Qwen)
Noometry Index39.431.9
Released—2024-09
WeightsOpenOpen
Context window—131K
Max output—8K
Input $ / M tokens—$1.40
Output $ / M tokens—$5.60
Results tracked1643

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 39.5 (#157), Qwen2.5 72B Instruct: 33.2 (#260)

Coding benchmarks
BenchmarkOlmo 3.1 32b InstructQwen2.5 72B Instruct
LMArena Coding13471292
WeirdML—16%
BigCodeBench Instruct—45.8%
BigCodeBench Complete—55.9%

Agentic & Tool Use Not comparable

Olmo 3.1 32b Instruct: —, Qwen2.5 72B Instruct: 22.1 (#133)

Agentic & Tool Use benchmarks
BenchmarkOlmo 3.1 32b InstructQwen2.5 72B Instruct
TheAgentCompany—5.7%
BALROG—16.2%
METR Time Horizons—35.8%

Reasoning Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 26.4 (#132), Qwen2.5 72B Instruct: 22.3 (#199)

Reasoning benchmarks
BenchmarkOlmo 3.1 32b InstructQwen2.5 72B Instruct
LMArena Hard Prompts13221271
DTBench—62.9%
LMCA—13.4%
BIG-Bench Hard—79.8%
Epoch Capabilities Index—129
ForecastBench—57.5
HellaSwag—84.8%
PIQA—82.6%
WinoGrande—82.3%

Math Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 36.3 (#167), Qwen2.5 72B Instruct: 19.3 (#287)

Math benchmarks
BenchmarkOlmo 3.1 32b InstructQwen2.5 72B Instruct
LMArena Math13051283
OTIS Mock AIME 2024-2025—8.1%
Omni-MATH—33%
MATH Level 5—63.2%

Knowledge Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 36.1 (#175), Qwen2.5 72B Instruct: 27.0 (#253)

Knowledge benchmarks
BenchmarkOlmo 3.1 32b InstructQwen2.5 72B Instruct
LMArena Expert13081245
GPQA Diamond—49.1%
MMLU-Pro—63.1%
Confabulations—19.1%
GPQA (HELM)—42.6%
ARC (AI2) Challenge—94.5%
MMLU—85.3%
TriviaQA—71.9%

Multilingual Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 42.6 (#191), Qwen2.5 72B Instruct: 41.0 (#213)

Multilingual benchmarks
BenchmarkOlmo 3.1 32b InstructQwen2.5 72B Instruct
LMArena Non-English12751252
LMArena Chinese13041272
LMArena French13281280
LMArena German12821234
LMArena Korean12061188
LMArena Russian12681264
LMArena Spanish13361256
LMArena Japanese—1180

Instruction Following Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 68.6 (#187), Qwen2.5 72B Instruct: 65.5 (#221)

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b InstructQwen2.5 72B Instruct
LMArena Instruction Following12991254
IFEval—80.6%

Long Context Too close to call

Olmo 3.1 32b Instruct: 39.9 (#166), Qwen2.5 72B Instruct: 38.9 (#188)

Long Context benchmarks
BenchmarkOlmo 3.1 32b InstructQwen2.5 72B Instruct
LMArena Longer Query13121282

Writing & Preference Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 50.2 (#185), Qwen2.5 72B Instruct: 46.7 (#215)

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b InstructQwen2.5 72B Instruct
LMArena Text13111269
LMArena Creative Writing12641221
LMArena Multi-Turn13091272
WildBench—80.2%

Frequently asked questions

Is Olmo 3.1 32b Instruct better than Qwen2.5 72B Instruct?

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 31.9 on the Noometry Index.

Is Olmo 3.1 32b Instruct or Qwen2.5 72B Instruct better for coding?

Olmo 3.1 32b Instruct scores higher on coding benchmarks: 39.5 versus 33.2 in the Noometry coding category.

How many benchmarks do Olmo 3.1 32b Instruct and Qwen2.5 72B Instruct share?

16 benchmarks have published results for both models. Olmo 3.1 32b Instruct has 16 scored results on Noometry and Qwen2.5 72B Instruct has 43.

Related comparisons

Go deeper