Model comparison

Olmo 3.1 32b Instruct vs Olmo 3.1 32b Think

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 37.9 on the Noometry Index.

Last verified . 15 shared benchmarks.

Summary

  • They share 15 benchmarks with published results for both. Olmo 3.1 32b Instruct scores higher in 7 categories and Olmo 3.1 32b Think in 1 category; 6 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Olmo 3.1 32b Instruct leads 42.6 to 38.1.

Side by side

Olmo 3.1 32b Instruct and Olmo 3.1 32b Think specifications
Olmo 3.1 32b InstructOlmo 3.1 32b Think
ProviderAllen Institute for AI (Ai2)Allen Institute for AI (Ai2)
Noometry Index39.437.9
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1615

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 39.5 (#157), Olmo 3.1 32b Think: 37.7 (#189)

Coding benchmarks
BenchmarkOlmo 3.1 32b InstructOlmo 3.1 32b Think
LMArena Coding13471291

Reasoning Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 26.4 (#132), Olmo 3.1 32b Think: 25.2 (#150)

Reasoning benchmarks
BenchmarkOlmo 3.1 32b InstructOlmo 3.1 32b Think
LMArena Hard Prompts13221272

Math Too close to call

Olmo 3.1 32b Instruct: 36.3 (#167), Olmo 3.1 32b Think: 36.3 (#168)

Math benchmarks
BenchmarkOlmo 3.1 32b InstructOlmo 3.1 32b Think
LMArena Math13051305

Knowledge Too close to call

Olmo 3.1 32b Instruct: 36.1 (#175), Olmo 3.1 32b Think: 35.7 (#181)

Knowledge benchmarks
BenchmarkOlmo 3.1 32b InstructOlmo 3.1 32b Think
LMArena Expert13081295

Multilingual Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 42.6 (#191), Olmo 3.1 32b Think: 38.1 (#231)

Multilingual benchmarks
BenchmarkOlmo 3.1 32b InstructOlmo 3.1 32b Think
LMArena Non-English12751209
LMArena Chinese13041242
LMArena French13281260
LMArena German12821262
LMArena Russian12681193
LMArena Spanish13361289
LMArena Korean1206—

Instruction Following Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 68.6 (#187), Olmo 3.1 32b Think: 65.6 (#218)

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b InstructOlmo 3.1 32b Think
LMArena Instruction Following12991247

Long Context Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 39.9 (#166), Olmo 3.1 32b Think: 38.6 (#195)

Long Context benchmarks
BenchmarkOlmo 3.1 32b InstructOlmo 3.1 32b Think
LMArena Longer Query13121272

Writing & Preference Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 50.2 (#185), Olmo 3.1 32b Think: 46.2 (#220)

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b InstructOlmo 3.1 32b Think
LMArena Text13111272
LMArena Creative Writing12641226
LMArena Multi-Turn13091252

Frequently asked questions

Is Olmo 3.1 32b Instruct better than Olmo 3.1 32b Think?

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 37.9 on the Noometry Index.

Is Olmo 3.1 32b Instruct or Olmo 3.1 32b Think better for coding?

Olmo 3.1 32b Instruct scores higher on coding benchmarks: 39.5 versus 37.7 in the Noometry coding category.

How many benchmarks do Olmo 3.1 32b Instruct and Olmo 3.1 32b Think share?

15 benchmarks have published results for both models. Olmo 3.1 32b Instruct has 16 scored results on Noometry and Olmo 3.1 32b Think has 15.

Related comparisons

Go deeper