Model comparison

Granite 4.1 8b vs Inkling

Inkling is the stronger model overall, scoring 44.1 to 37.4 on the Noometry Index.

Last verified . 13 shared benchmarks.

Granite 4.1 8b IBM

37.4

Rank #205 Confirmed

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Granite 4.1 8b scores higher in 1 category and Inkling in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Inkling leads 55.1 to 36.1.

Side by side

Granite 4.1 8b and Inkling specifications
Granite 4.1 8bInkling
ProviderIBMThinking Machines Lab
Noometry Index37.444.1
Released—2026-07-15
WeightsOpenOpen
Context window—66K
Max output—66K
Input $ / M tokens—$1.87
Output $ / M tokens—$4.68
Results tracked1341

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Inkling leads

Granite 4.1 8b: 30.1 (#297), Inkling: 34.5 (#234)

Coding benchmarks
BenchmarkGranite 4.1 8bInkling
LMArena WebDev11921413
LMArena Coding13121464
FrontierCode—14%
FrontierSWE—4.1%
SciCode—47%
WeirdML—32.3%
ALE-Bench—946

Agentic & Tool Use Not comparable

Granite 4.1 8b: —, Inkling: 29.6 (#85)

Agentic & Tool Use benchmarks
BenchmarkGranite 4.1 8bInkling
APEX-Agents—33.8%
τ²-bench Banking—25%

Reasoning Inkling leads

Granite 4.1 8b: 25.7 (#143), Inkling: 40.4 (#56)

Reasoning benchmarks
BenchmarkGranite 4.1 8bInkling
LMArena Hard Prompts12931451
ARC-AGI-2—36.5%
SimpleBench—50%
ARC-AGI-1—79.5%
CritPt—5.4%
Chess Puzzles—21%
DTBench—87.5%
LMCA—37.6%
Epoch Capabilities Index—148.54

Math Granite 4.1 8b leads

Granite 4.1 8b: 36.4 (#166), Inkling: 31.3 (#225)

Math benchmarks
BenchmarkGranite 4.1 8bInkling
LMArena Math13121479
FrontierMath (Tiers 1-3)—33.3%
FrontierMath Tier 4—4.9%
OTIS Mock AIME 2024-2025—88.9%
ProofBench—0%

Knowledge Inkling leads

Granite 4.1 8b: 36.1 (#174), Inkling: 55.1 (#49)

Knowledge benchmarks
BenchmarkGranite 4.1 8bInkling
LMArena Expert13091465
GPQA Diamond—88.3%
SimpleQA Verified—40.3%

Multilingual Inkling leads

Granite 4.1 8b: 41.7 (#204), Inkling: 54.0 (#52)

Multilingual benchmarks
BenchmarkGranite 4.1 8bInkling
LMArena Non-English12611434
LMArena Chinese13371490
LMArena Russian12401429
LMArena French—1458
LMArena German—1446
LMArena Japanese—1429
LMArena Korean—1404
LMArena Spanish—1448

Instruction Following Inkling leads

Granite 4.1 8b: 66.9 (#203), Inkling: 75.1 (#71)

Instruction Following benchmarks
BenchmarkGranite 4.1 8bInkling
LMArena Instruction Following12691426

Long Context Inkling leads

Granite 4.1 8b: 38.7 (#193), Inkling: 43.8 (#86)

Long Context benchmarks
BenchmarkGranite 4.1 8bInkling
LMArena Longer Query12751434

Writing & Preference Inkling leads

Granite 4.1 8b: 48.1 (#204), Inkling: 65.2 (#51)

Writing & Preference benchmarks
BenchmarkGranite 4.1 8bInkling
LMArena Text12901441
LMArena Creative Writing12511387
LMArena Multi-Turn12681436
EQ-Bench Creative Writing—1611
EQ-Bench 4—1226

Frequently asked questions

Is Granite 4.1 8b better than Inkling?

Inkling is the stronger model overall, scoring 44.1 to 37.4 on the Noometry Index.

Is Granite 4.1 8b or Inkling better for coding?

Inkling scores higher on coding benchmarks: 34.5 versus 30.1 in the Noometry coding category.

How many benchmarks do Granite 4.1 8b and Inkling share?

13 benchmarks have published results for both models. Granite 4.1 8b has 13 scored results on Noometry and Inkling has 41.

Related comparisons

Go deeper