Model comparison

Granite 4.2 30b vs Mercury 2

Granite 4.2 30b is the stronger model overall, scoring 41.8 to 39.1 on the Noometry Index.

Last verified . 11 shared benchmarks.

Granite 4.2 30b IBM

41.8

Rank #130 Confirmed

Mercury 2 Inception

39.1

Rank #175 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Granite 4.2 30b scores higher in 6 categories and Mercury 2 in 1 category; 3 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Granite 4.2 30b leads 41.0 to 33.5.
  • Granite 4.2 30b has downloadable open weights; the other is API-only.

Side by side

Granite 4.2 30b and Mercury 2 specifications
Granite 4.2 30bMercury 2
ProviderIBMInception
Noometry Index41.839.1
Released—2026-02-20
WeightsOpenProprietary
Context window—128K
Max output—50K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.75
Results tracked1117

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Granite 4.2 30b leads

Granite 4.2 30b: 41.0 (#126), Mercury 2: 33.5 (#255)

Coding benchmarks
BenchmarkGranite 4.2 30bMercury 2
LMArena Coding13961391
LMArena WebDev—1171
SciCode—38.7%
WeirdML—43.2%
ALE-Bench—785.58

Reasoning Granite 4.2 30b leads

Granite 4.2 30b: 27.8 (#112), Mercury 2: 23.8 (#170)

Reasoning benchmarks
BenchmarkGranite 4.2 30bMercury 2
LMArena Hard Prompts13741362
CritPt—0.8%

Knowledge Granite 4.2 30b leads

Granite 4.2 30b: 39.1 (#138), Mercury 2: 36.2 (#172)

Knowledge benchmarks
BenchmarkGranite 4.2 30bMercury 2
LMArena Expert14061358
Vectara Hallucination Rate—12.3%

Multilingual Too close to call

Granite 4.2 30b: 47.3 (#151), Mercury 2: 46.6 (#157)

Multilingual benchmarks
BenchmarkGranite 4.2 30bMercury 2
LMArena Non-English13401331
LMArena Chinese14141417
LMArena Russian13431304

Instruction Following Too close to call

Granite 4.2 30b: 71.2 (#155), Mercury 2: 70.2 (#165)

Instruction Following benchmarks
BenchmarkGranite 4.2 30bMercury 2
LMArena Instruction Following13471329

Long Context Too close to call

Granite 4.2 30b: 41.4 (#140), Mercury 2: 40.5 (#154)

Long Context benchmarks
BenchmarkGranite 4.2 30bMercury 2
LMArena Longer Query13591330

Writing & Preference Too close to call

Granite 4.2 30b: 53.8 (#156), Mercury 2: 53.8 (#155)

Writing & Preference benchmarks
BenchmarkGranite 4.2 30bMercury 2
LMArena Text13611355
LMArena Creative Writing12881289
LMArena Multi-Turn13391358

Frequently asked questions

Is Granite 4.2 30b better than Mercury 2?

Granite 4.2 30b is the stronger model overall, scoring 41.8 to 39.1 on the Noometry Index.

Is Granite 4.2 30b or Mercury 2 better for coding?

Granite 4.2 30b scores higher on coding benchmarks: 41.0 versus 33.5 in the Noometry coding category.

How many benchmarks do Granite 4.2 30b and Mercury 2 share?

11 benchmarks have published results for both models. Granite 4.2 30b has 11 scored results on Noometry and Mercury 2 has 17.

Related comparisons

Go deeper