Model comparison

Mercury 2 vs Olmo 3.1 32b Instruct

Mercury 2 and Olmo 3.1 32b Instruct score almost the same on the Noometry Index (39.1 vs 39.4), so choose on price, context window or the category you care about most.

Last verified . 11 shared benchmarks.

Mercury 2 Inception

39.1

Rank #175 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Mercury 2 scores higher in 5 categories and Olmo 3.1 32b Instruct in 2 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Olmo 3.1 32b Instruct leads 39.5 to 33.5.
  • Olmo 3.1 32b Instruct has downloadable open weights; the other is API-only.

Side by side

Mercury 2 and Olmo 3.1 32b Instruct specifications
Mercury 2Olmo 3.1 32b Instruct
ProviderInceptionAllen Institute for AI (Ai2)
Noometry Index39.139.4
Released2026-02-20—
WeightsProprietaryOpen
Context window128K—
Max output50K—
Input $ / M tokens$0.25—
Output $ / M tokens$0.75—
Results tracked1716

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Instruct leads

Mercury 2: 33.5 (#255), Olmo 3.1 32b Instruct: 39.5 (#157)

Coding benchmarks
BenchmarkMercury 2Olmo 3.1 32b Instruct
LMArena Coding13911347
LMArena WebDev1171—
SciCode38.7%—
WeirdML43.2%—
ALE-Bench785.58—

Reasoning Olmo 3.1 32b Instruct leads

Mercury 2: 23.8 (#170), Olmo 3.1 32b Instruct: 26.4 (#132)

Reasoning benchmarks
BenchmarkMercury 2Olmo 3.1 32b Instruct
LMArena Hard Prompts13621322
CritPt0.8%—

Math Not comparable

Mercury 2: —, Olmo 3.1 32b Instruct: 36.3 (#167)

Math benchmarks
BenchmarkMercury 2Olmo 3.1 32b Instruct
LMArena Math—1305

Knowledge Too close to call

Mercury 2: 36.2 (#172), Olmo 3.1 32b Instruct: 36.1 (#175)

Knowledge benchmarks
BenchmarkMercury 2Olmo 3.1 32b Instruct
LMArena Expert13581308
Vectara Hallucination Rate12.3%—

Multilingual Mercury 2 leads

Mercury 2: 46.6 (#157), Olmo 3.1 32b Instruct: 42.6 (#191)

Multilingual benchmarks
BenchmarkMercury 2Olmo 3.1 32b Instruct
LMArena Non-English13311275
LMArena Chinese14171304
LMArena Russian13041268
LMArena French—1328
LMArena German—1282
LMArena Korean—1206
LMArena Spanish—1336

Instruction Following Mercury 2 leads

Mercury 2: 70.2 (#165), Olmo 3.1 32b Instruct: 68.6 (#187)

Instruction Following benchmarks
BenchmarkMercury 2Olmo 3.1 32b Instruct
LMArena Instruction Following13291299

Long Context Too close to call

Mercury 2: 40.5 (#154), Olmo 3.1 32b Instruct: 39.9 (#166)

Long Context benchmarks
BenchmarkMercury 2Olmo 3.1 32b Instruct
LMArena Longer Query13301312

Writing & Preference Mercury 2 leads

Mercury 2: 53.8 (#155), Olmo 3.1 32b Instruct: 50.2 (#185)

Writing & Preference benchmarks
BenchmarkMercury 2Olmo 3.1 32b Instruct
LMArena Text13551311
LMArena Creative Writing12891264
LMArena Multi-Turn13581309

Frequently asked questions

Is Mercury 2 better than Olmo 3.1 32b Instruct?

Mercury 2 and Olmo 3.1 32b Instruct score almost the same on the Noometry Index (39.1 vs 39.4), so choose on price, context window or the category you care about most.

Is Mercury 2 or Olmo 3.1 32b Instruct better for coding?

Olmo 3.1 32b Instruct scores higher on coding benchmarks: 39.5 versus 33.5 in the Noometry coding category.

How many benchmarks do Mercury 2 and Olmo 3.1 32b Instruct share?

11 benchmarks have published results for both models. Mercury 2 has 17 scored results on Noometry and Olmo 3.1 32b Instruct has 16.

Related comparisons

Go deeper