Model comparison

Mercury vs Olmo 3.1 32b Think

Mercury and Olmo 3.1 32b Think score almost the same on the Noometry Index (37.6 vs 37.9), so choose on price, context window or the category you care about most.

Last verified . 8 shared benchmarks.

Mercury Inception

37.6

Rank #199 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Mercury scores higher in 2 categories and Olmo 3.1 32b Think in 4 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Olmo 3.1 32b Think leads 25.2 to 17.5.
  • Olmo 3.1 32b Think has downloadable open weights; the other is API-only.

Side by side

Mercury and Olmo 3.1 32b Think specifications
MercuryOlmo 3.1 32b Think
ProviderInceptionAllen Institute for AI (Ai2)
Noometry Index37.637.9
Released——
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked915

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Mercury: 38.7 (#170), Olmo 3.1 32b Think: 37.7 (#189)

Coding benchmarks
BenchmarkMercuryOlmo 3.1 32b Think
LMArena Coding13221291

Reasoning Olmo 3.1 32b Think leads

Mercury: 17.5 (#293), Olmo 3.1 32b Think: 25.2 (#150)

Reasoning benchmarks
BenchmarkMercuryOlmo 3.1 32b Think
LMArena Hard Prompts12851272
Kagi LLM Benchmark21.6%—

Math Not comparable

Mercury: —, Olmo 3.1 32b Think: 36.3 (#168)

Math benchmarks
BenchmarkMercuryOlmo 3.1 32b Think
LMArena Math—1305

Knowledge Not comparable

Mercury: —, Olmo 3.1 32b Think: 35.7 (#181)

Knowledge benchmarks
BenchmarkMercuryOlmo 3.1 32b Think
LMArena Expert—1295

Multilingual Mercury leads

Mercury: 41.6 (#206), Olmo 3.1 32b Think: 38.1 (#231)

Multilingual benchmarks
BenchmarkMercuryOlmo 3.1 32b Think
LMArena Non-English12601209
LMArena Chinese—1242
LMArena French—1260
LMArena German—1262
LMArena Russian—1193
LMArena Spanish—1289

Instruction Following Too close to call

Mercury: 65.2 (#224), Olmo 3.1 32b Think: 65.6 (#218)

Instruction Following benchmarks
BenchmarkMercuryOlmo 3.1 32b Think
LMArena Instruction Following12391247

Long Context Too close to call

Mercury: 38.4 (#198), Olmo 3.1 32b Think: 38.6 (#195)

Long Context benchmarks
BenchmarkMercuryOlmo 3.1 32b Think
LMArena Longer Query12661272

Writing & Preference Too close to call

Mercury: 46.2 (#221), Olmo 3.1 32b Think: 46.2 (#220)

Writing & Preference benchmarks
BenchmarkMercuryOlmo 3.1 32b Think
LMArena Text12821272
LMArena Creative Writing11911226
LMArena Multi-Turn12821252

Frequently asked questions

Is Mercury better than Olmo 3.1 32b Think?

Mercury and Olmo 3.1 32b Think score almost the same on the Noometry Index (37.6 vs 37.9), so choose on price, context window or the category you care about most.

Is Mercury or Olmo 3.1 32b Think better for coding?

They score almost the same on coding (38.7 vs 37.7); test both on your own repository before choosing.

How many benchmarks do Mercury and Olmo 3.1 32b Think share?

8 benchmarks have published results for both models. Mercury has 9 scored results on Noometry and Olmo 3.1 32b Think has 15.

Related comparisons

Go deeper