Model comparison

Granite 4.0 Micro vs Grok 4.1

Grok 4.1 is the stronger model overall, scoring 41.5 to 29.0 on the Noometry Index.

Last verified . 0 shared benchmarks.

Granite 4.0 Micro IBM

29.0

Rank #318 Confirmed

Grok 4.1 xAI

41.5

Rank #134 Confirmed

Summary

  • The widest gap is in knowledge, where Grok 4.1 leads 39.5 to 9.9.
  • Granite 4.0 Micro has downloadable open weights; the other is API-only.

Side by side

Granite 4.0 Micro and Grok 4.1 specifications
Granite 4.0 MicroGrok 4.1
ProviderIBMxAI
Noometry Index29.041.5
Released2025-10-022025-11-17
WeightsOpenProprietary
Context window131K—
Max output118K—
Input $ / M tokens$0.017—
Output $ / M tokens$0.11—
Results tracked819

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Granite 4.0 Micro: —, Grok 4.1: 33.7 (#253)

Coding benchmarks
BenchmarkGranite 4.0 MicroGrok 4.1
LMArena WebDev—1214
LMArena Coding—1445

Agentic & Tool Use Not comparable

Granite 4.0 Micro: —, Grok 4.1: 34.1 (#49)

Agentic & Tool Use benchmarks
BenchmarkGranite 4.0 MicroGrok 4.1
Cybench—39%

Reasoning Grok 4.1 leads

Granite 4.0 Micro: 19.2 (#265), Grok 4.1: 29.5 (#91)

Reasoning benchmarks
BenchmarkGranite 4.0 MicroGrok 4.1
Chess Puzzles0%—
LMArena Hard Prompts—1435

Math Grok 4.1 leads

Granite 4.0 Micro: 12.0 (#307), Grok 4.1: 38.9 (#120)

Math benchmarks
BenchmarkGranite 4.0 MicroGrok 4.1
OTIS Mock AIME 2024-20252.8%—
Omni-MATH20.9%—
LMArena Math—1422

Knowledge Grok 4.1 leads

Granite 4.0 Micro: 9.9 (#304), Grok 4.1: 39.5 (#133)

Knowledge benchmarks
BenchmarkGranite 4.0 MicroGrok 4.1
GPQA Diamond28.3%—
MMLU-Pro39.5%—
GPQA (HELM)30.7%—
LMArena Expert—1417

Multilingual Not comparable

Granite 4.0 Micro: —, Grok 4.1: 53.4 (#68)

Multilingual benchmarks
BenchmarkGranite 4.0 MicroGrok 4.1
LMArena Non-English—1425
LMArena Chinese—1465
LMArena French—1448
LMArena German—1446
LMArena Japanese—1397
LMArena Korean—1407
LMArena Russian—1434
LMArena Spanish—1438

Instruction Following Grok 4.1 leads

Granite 4.0 Micro: 69.9 (#169), Grok 4.1: 73.8 (#111)

Instruction Following benchmarks
BenchmarkGranite 4.0 MicroGrok 4.1
IFEval84.9%—
LMArena Instruction Following—1400

Long Context Not comparable

Granite 4.0 Micro: —, Grok 4.1: 43.2 (#100)

Long Context benchmarks
BenchmarkGranite 4.0 MicroGrok 4.1
LMArena Longer Query—1416

Writing & Preference Grok 4.1 leads

Granite 4.0 Micro: 46.7 (#216), Grok 4.1: 62.4 (#75)

Writing & Preference benchmarks
BenchmarkGranite 4.0 MicroGrok 4.1
LMArena Text—1437
LMArena Creative Writing—1411
WildBench67%—
LMArena Multi-Turn—1437

Frequently asked questions

Is Granite 4.0 Micro better than Grok 4.1?

Grok 4.1 is the stronger model overall, scoring 41.5 to 29.0 on the Noometry Index.

How many benchmarks do Granite 4.0 Micro and Grok 4.1 share?

0 benchmarks have published results for both models. Granite 4.0 Micro has 8 scored results on Noometry and Grok 4.1 has 19.

Related comparisons

Go deeper