Model comparison

Olmo 3.1 32b Instruct vs Olmo 7b Instruct

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 30.3 on the Noometry Index.

Last verified . 10 shared benchmarks.

Summary

  • They share 10 benchmarks with published results for both. Olmo 3.1 32b Instruct scores higher in 6 categories and Olmo 7b Instruct in 0 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Olmo 3.1 32b Instruct leads 50.2 to 25.8.

Side by side

Olmo 3.1 32b Instruct and Olmo 7b Instruct specifications
Olmo 3.1 32b InstructOlmo 7b Instruct
ProviderAllen Institute for AI (Ai2)Allen Institute for AI (Ai2)
Noometry Index39.430.3
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1610

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 39.5 (#157), Olmo 7b Instruct: 29.6 (#303)

Coding benchmarks
BenchmarkOlmo 3.1 32b InstructOlmo 7b Instruct
LMArena Coding13471016

Reasoning Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 26.4 (#132), Olmo 7b Instruct: 18.8 (#274)

Reasoning benchmarks
BenchmarkOlmo 3.1 32b InstructOlmo 7b Instruct
LMArena Hard Prompts1322993

Math Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 36.3 (#167), Olmo 7b Instruct: 30.2 (#237)

Math benchmarks
BenchmarkOlmo 3.1 32b InstructOlmo 7b Instruct
LMArena Math13051018

Knowledge Not comparable

Olmo 3.1 32b Instruct: 36.1 (#175), Olmo 7b Instruct: —

Knowledge benchmarks
BenchmarkOlmo 3.1 32b InstructOlmo 7b Instruct
LMArena Expert1308—

Multilingual Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 42.6 (#191), Olmo 7b Instruct: 24.0 (#291)

Multilingual benchmarks
BenchmarkOlmo 3.1 32b InstructOlmo 7b Instruct
LMArena Non-English1275977
LMArena Chinese13041014
LMArena Russian1268947
LMArena French1328—
LMArena German1282—
LMArena Korean1206—
LMArena Spanish1336—

Instruction Following Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 68.6 (#187), Olmo 7b Instruct: 49.0 (#301)

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b InstructOlmo 7b Instruct
LMArena Instruction Following1299978

Long Context Not comparable

Olmo 3.1 32b Instruct: 39.9 (#166), Olmo 7b Instruct: —

Long Context benchmarks
BenchmarkOlmo 3.1 32b InstructOlmo 7b Instruct
LMArena Longer Query1312—

Writing & Preference Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 50.2 (#185), Olmo 7b Instruct: 25.8 (#303)

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b InstructOlmo 7b Instruct
LMArena Text13111032
LMArena Creative Writing1264990
LMArena Multi-Turn13091007

Frequently asked questions

Is Olmo 3.1 32b Instruct better than Olmo 7b Instruct?

Olmo 3.1 32b Instruct is the stronger model overall, scoring 39.4 to 30.3 on the Noometry Index.

Is Olmo 3.1 32b Instruct or Olmo 7b Instruct better for coding?

Olmo 3.1 32b Instruct scores higher on coding benchmarks: 39.5 versus 29.6 in the Noometry coding category.

How many benchmarks do Olmo 3.1 32b Instruct and Olmo 7b Instruct share?

10 benchmarks have published results for both models. Olmo 3.1 32b Instruct has 16 scored results on Noometry and Olmo 7b Instruct has 10.

Related comparisons

Go deeper