Model comparison

Olmo 3.1 32b Think vs Sonar

Olmo 3.1 32b Think and Sonar score almost the same on the Noometry Index (37.9 vs 38.5), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

Sonar Perplexity

38.5

Rank #187 Confirmed

Summary

  • The widest gap is in writing & preference, where Sonar leads 52.6 to 46.2.
  • Olmo 3.1 32b Think has downloadable open weights; the other is API-only.

Side by side

Olmo 3.1 32b Think and Sonar specifications
Olmo 3.1 32b ThinkSonar
ProviderAllen Institute for AI (Ai2)Perplexity
Noometry Index37.938.5
Released—2024-01-01
WeightsOpenProprietary
Context window—128K
Max output—4K
Input $ / M tokens—$1
Output $ / M tokens—$1
Results tracked157

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Think leads

Olmo 3.1 32b Think: 37.7 (#189), Sonar: 35.7 (#221)

Coding benchmarks
BenchmarkOlmo 3.1 32b ThinkSonar
LiveBench Coding—35.1%
LMArena Coding1291—

Reasoning Olmo 3.1 32b Think leads

Olmo 3.1 32b Think: 25.2 (#150), Sonar: 21.1 (#227)

Reasoning benchmarks
BenchmarkOlmo 3.1 32b ThinkSonar
LiveBench Reasoning—46.3%
LMArena Hard Prompts1272—
LiveBench Data Analysis—37.9%
LiveBench—46.9%

Math Olmo 3.1 32b Think leads

Olmo 3.1 32b Think: 36.3 (#168), Sonar: 33.7 (#200)

Math benchmarks
BenchmarkOlmo 3.1 32b ThinkSonar
LiveBench Math—41.6%
LMArena Math1305—

Knowledge Not comparable

Olmo 3.1 32b Think: 35.7 (#181), Sonar: —

Knowledge benchmarks
BenchmarkOlmo 3.1 32b ThinkSonar
LMArena Expert1295—

Multilingual Not comparable

Olmo 3.1 32b Think: 38.1 (#231), Sonar: —

Multilingual benchmarks
BenchmarkOlmo 3.1 32b ThinkSonar
LMArena Non-English1209—
LMArena Chinese1242—
LMArena French1260—
LMArena German1262—
LMArena Russian1193—
LMArena Spanish1289—

Instruction Following Sonar leads

Olmo 3.1 32b Think: 65.6 (#218), Sonar: 71.4 (#150)

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b ThinkSonar
LiveBench Instruction Following—76.2%
LMArena Instruction Following1247—

Long Context Not comparable

Olmo 3.1 32b Think: 38.6 (#195), Sonar: —

Long Context benchmarks
BenchmarkOlmo 3.1 32b ThinkSonar
LMArena Longer Query1272—

Writing & Preference Sonar leads

Olmo 3.1 32b Think: 46.2 (#220), Sonar: 52.6 (#167)

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b ThinkSonar
LMArena Text1272—
LMArena Creative Writing1226—
LMArena Multi-Turn1252—
LiveBench Language—44.1%

Frequently asked questions

Is Olmo 3.1 32b Think better than Sonar?

Olmo 3.1 32b Think and Sonar score almost the same on the Noometry Index (37.9 vs 38.5), so choose on price, context window or the category you care about most.

Is Olmo 3.1 32b Think or Sonar better for coding?

Olmo 3.1 32b Think scores higher on coding benchmarks: 37.7 versus 35.7 in the Noometry coding category.

How many benchmarks do Olmo 3.1 32b Think and Sonar share?

0 benchmarks have published results for both models. Olmo 3.1 32b Think has 15 scored results on Noometry and Sonar has 7.

Related comparisons

Go deeper