Model comparison

Granite 4.2 8B vs Olmo 3.1 32b Think

Granite 4.2 8B is the stronger model overall, scoring 40.5 to 37.9 on the Noometry Index.

Last verified . 11 shared benchmarks.

Granite 4.2 8B IBM

40.5

Rank #148 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Granite 4.2 8B scores higher in 7 categories and Olmo 3.1 32b Think in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Granite 4.2 8B leads 44.5 to 38.1.

Side by side

Granite 4.2 8B and Olmo 3.1 32b Think specifications
Granite 4.2 8BOlmo 3.1 32b Think
ProviderIBMAllen Institute for AI (Ai2)
Noometry Index40.537.9
Released——
WeightsOpenOpen
Context window131K—
Max output118K—
Input $ / M tokens$0.06—
Output $ / M tokens$0.25—
Results tracked1115

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Granite 4.2 8B leads

Granite 4.2 8B: 40.5 (#137), Olmo 3.1 32b Think: 37.7 (#189)

Coding benchmarks
BenchmarkGranite 4.2 8BOlmo 3.1 32b Think
LMArena Coding13801291

Reasoning Granite 4.2 8B leads

Granite 4.2 8B: 26.6 (#131), Olmo 3.1 32b Think: 25.2 (#150)

Reasoning benchmarks
BenchmarkGranite 4.2 8BOlmo 3.1 32b Think
LMArena Hard Prompts13291272

Math Not comparable

Granite 4.2 8B: —, Olmo 3.1 32b Think: 36.3 (#168)

Math benchmarks
BenchmarkGranite 4.2 8BOlmo 3.1 32b Think
LMArena Math—1305

Knowledge Granite 4.2 8B leads

Granite 4.2 8B: 38.4 (#145), Olmo 3.1 32b Think: 35.7 (#181)

Knowledge benchmarks
BenchmarkGranite 4.2 8BOlmo 3.1 32b Think
LMArena Expert13841295

Multilingual Granite 4.2 8B leads

Granite 4.2 8B: 44.5 (#178), Olmo 3.1 32b Think: 38.1 (#231)

Multilingual benchmarks
BenchmarkGranite 4.2 8BOlmo 3.1 32b Think
LMArena Non-English13021209
LMArena Chinese13661242
LMArena Russian12851193
LMArena French—1260
LMArena German—1262
LMArena Spanish—1289

Instruction Following Granite 4.2 8B leads

Granite 4.2 8B: 68.7 (#184), Olmo 3.1 32b Think: 65.6 (#218)

Instruction Following benchmarks
BenchmarkGranite 4.2 8BOlmo 3.1 32b Think
LMArena Instruction Following13011247

Long Context Granite 4.2 8B leads

Granite 4.2 8B: 40.3 (#159), Olmo 3.1 32b Think: 38.6 (#195)

Long Context benchmarks
BenchmarkGranite 4.2 8BOlmo 3.1 32b Think
LMArena Longer Query13241272

Writing & Preference Granite 4.2 8B leads

Granite 4.2 8B: 49.6 (#189), Olmo 3.1 32b Think: 46.2 (#220)

Writing & Preference benchmarks
BenchmarkGranite 4.2 8BOlmo 3.1 32b Think
LMArena Text13201272
LMArena Creative Writing12361226
LMArena Multi-Turn13011252

Frequently asked questions

Is Granite 4.2 8B better than Olmo 3.1 32b Think?

Granite 4.2 8B is the stronger model overall, scoring 40.5 to 37.9 on the Noometry Index.

Is Granite 4.2 8B or Olmo 3.1 32b Think better for coding?

Granite 4.2 8B scores higher on coding benchmarks: 40.5 versus 37.7 in the Noometry coding category.

How many benchmarks do Granite 4.2 8B and Olmo 3.1 32b Think share?

11 benchmarks have published results for both models. Granite 4.2 8B has 11 scored results on Noometry and Olmo 3.1 32b Think has 15.

Related comparisons

Go deeper