Model comparison
Granite 4.0 Micro vs Mercury 2
Mercury 2 is the stronger model overall, scoring 39.1 to 29.0 on the Noometry Index. Granite 4.0 Micro costs 9.2× less per token, which makes it the better buy when Mercury 2's lead doesn't matter for your workload.
Last verified . 0 shared benchmarks.
Summary
- The widest gap is in knowledge, where Mercury 2 leads 36.2 to 9.9.
- Granite 4.0 Micro is cheaper at $0.017 / $0.11 per million input/output tokens, against $0.25 / $0.75 for Mercury 2.
- Granite 4.0 Micro accepts more context: 131K tokens versus 128K.
- Granite 4.0 Micro has downloadable open weights; the other is API-only.
Side by side
| Granite 4.0 Micro | Mercury 2 | |
|---|---|---|
| Provider | IBM | Inception |
| Noometry Index | 29.0 | 39.1 |
| Released | 2025-10-02 | 2026-02-20 |
| Weights | Open | Proprietary |
| Context window | 131K | 128K |
| Max output | 118K | 50K |
| Input $ / M tokens | $0.017 | $0.25 |
| Output $ / M tokens | $0.11 | $0.75 |
| Results tracked | 8 | 17 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Not comparable
Granite 4.0 Micro: —, Mercury 2: 33.5 (#255)
| Benchmark | Granite 4.0 Micro | Mercury 2 |
|---|---|---|
| LMArena WebDev | — | 1171 |
| SciCode | — | 38.7% |
| WeirdML | — | 43.2% |
| LMArena Coding | — | 1391 |
| ALE-Bench | — | 785.58 |
Reasoning Mercury 2 leads
Granite 4.0 Micro: 19.2 (#265), Mercury 2: 23.8 (#170)
| Benchmark | Granite 4.0 Micro | Mercury 2 |
|---|---|---|
| CritPt | — | 0.8% |
| Chess Puzzles | 0% | — |
| LMArena Hard Prompts | — | 1362 |
Math Not comparable
Granite 4.0 Micro: 12.0 (#307), Mercury 2: —
| Benchmark | Granite 4.0 Micro | Mercury 2 |
|---|---|---|
| OTIS Mock AIME 2024-2025 | 2.8% | — |
| Omni-MATH | 20.9% | — |
Knowledge Mercury 2 leads
Granite 4.0 Micro: 9.9 (#304), Mercury 2: 36.2 (#172)
| Benchmark | Granite 4.0 Micro | Mercury 2 |
|---|---|---|
| GPQA Diamond | 28.3% | — |
| MMLU-Pro | 39.5% | — |
| Vectara Hallucination Rate | — | 12.3% |
| GPQA (HELM) | 30.7% | — |
| LMArena Expert | — | 1358 |
Multilingual Not comparable
Granite 4.0 Micro: —, Mercury 2: 46.6 (#157)
| Benchmark | Granite 4.0 Micro | Mercury 2 |
|---|---|---|
| LMArena Non-English | — | 1331 |
| LMArena Chinese | — | 1417 |
| LMArena Russian | — | 1304 |
Instruction Following Too close to call
Granite 4.0 Micro: 69.9 (#169), Mercury 2: 70.2 (#165)
| Benchmark | Granite 4.0 Micro | Mercury 2 |
|---|---|---|
| IFEval | 84.9% | — |
| LMArena Instruction Following | — | 1329 |
Long Context Not comparable
Granite 4.0 Micro: —, Mercury 2: 40.5 (#154)
| Benchmark | Granite 4.0 Micro | Mercury 2 |
|---|---|---|
| LMArena Longer Query | — | 1330 |
Writing & Preference Mercury 2 leads
Granite 4.0 Micro: 46.7 (#216), Mercury 2: 53.8 (#155)
| Benchmark | Granite 4.0 Micro | Mercury 2 |
|---|---|---|
| LMArena Text | — | 1355 |
| LMArena Creative Writing | — | 1289 |
| WildBench | 67% | — |
| LMArena Multi-Turn | — | 1358 |
Frequently asked questions
Is Granite 4.0 Micro better than Mercury 2?
Mercury 2 is the stronger model overall, scoring 39.1 to 29.0 on the Noometry Index. Granite 4.0 Micro costs 9.2× less per token, which makes it the better buy when Mercury 2's lead doesn't matter for your workload.
Which is cheaper, Granite 4.0 Micro or Mercury 2?
Granite 4.0 Micro is cheaper. It lists at $0.017 per million input tokens and $0.11 per million output tokens; Mercury 2 lists at $0.25 and $0.75.
Which has the bigger context window?
Granite 4.0 Micro does, with 131K tokens against 128K.
How many benchmarks do Granite 4.0 Micro and Mercury 2 share?
0 benchmarks have published results for both models. Granite 4.0 Micro has 8 scored results on Noometry and Mercury 2 has 17.