Model comparison

Olmo 3 32b Think vs Olmo 7b Instruct

Olmo 3 32b Think is the stronger model overall, scoring 38.7 to 30.3 on the Noometry Index.

Last verified . 10 shared benchmarks.

Summary

  • They share 10 benchmarks with published results for both. Olmo 3 32b Think scores higher in 6 categories and Olmo 7b Instruct in 0 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Olmo 3 32b Think leads 49.1 to 25.8.

Side by side

Olmo 3 32b Think and Olmo 7b Instruct specifications
Olmo 3 32b ThinkOlmo 7b Instruct
ProviderAllen Institute for AI (Ai2)Allen Institute for AI (Ai2)
Noometry Index38.730.3
Released——
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1410

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3 32b Think leads

Olmo 3 32b Think: 38.6 (#172), Olmo 7b Instruct: 29.6 (#303)

Coding benchmarks
BenchmarkOlmo 3 32b ThinkOlmo 7b Instruct
LMArena Coding13191016

Reasoning Olmo 3 32b Think leads

Olmo 3 32b Think: 25.9 (#140), Olmo 7b Instruct: 18.8 (#274)

Reasoning benchmarks
BenchmarkOlmo 3 32b ThinkOlmo 7b Instruct
LMArena Hard Prompts1302993

Math Olmo 3 32b Think leads

Olmo 3 32b Think: 36.5 (#165), Olmo 7b Instruct: 30.2 (#237)

Math benchmarks
BenchmarkOlmo 3 32b ThinkOlmo 7b Instruct
LMArena Math13161018

Knowledge Not comparable

Olmo 3 32b Think: 35.0 (#190), Olmo 7b Instruct: —

Knowledge benchmarks
BenchmarkOlmo 3 32b ThinkOlmo 7b Instruct
LMArena Expert1273—

Multilingual Olmo 3 32b Think leads

Olmo 3 32b Think: 41.2 (#210), Olmo 7b Instruct: 24.0 (#291)

Multilingual benchmarks
BenchmarkOlmo 3 32b ThinkOlmo 7b Instruct
LMArena Non-English1255977
LMArena Chinese13001014
LMArena Russian1254947
LMArena French1291—
LMArena German1290—

Instruction Following Olmo 3 32b Think leads

Olmo 3 32b Think: 67.2 (#198), Olmo 7b Instruct: 49.0 (#301)

Instruction Following benchmarks
BenchmarkOlmo 3 32b ThinkOlmo 7b Instruct
LMArena Instruction Following1275978

Long Context Not comparable

Olmo 3 32b Think: 39.4 (#182), Olmo 7b Instruct: —

Long Context benchmarks
BenchmarkOlmo 3 32b ThinkOlmo 7b Instruct
LMArena Longer Query1296—

Writing & Preference Olmo 3 32b Think leads

Olmo 3 32b Think: 49.1 (#193), Olmo 7b Instruct: 25.8 (#303)

Writing & Preference benchmarks
BenchmarkOlmo 3 32b ThinkOlmo 7b Instruct
LMArena Text13001032
LMArena Creative Writing1256990
LMArena Multi-Turn12901007

Frequently asked questions

Is Olmo 3 32b Think better than Olmo 7b Instruct?

Olmo 3 32b Think is the stronger model overall, scoring 38.7 to 30.3 on the Noometry Index.

Is Olmo 3 32b Think or Olmo 7b Instruct better for coding?

Olmo 3 32b Think scores higher on coding benchmarks: 38.6 versus 29.6 in the Noometry coding category.

How many benchmarks do Olmo 3 32b Think and Olmo 7b Instruct share?

10 benchmarks have published results for both models. Olmo 3 32b Think has 14 scored results on Noometry and Olmo 7b Instruct has 10.

Related comparisons

Go deeper