Model comparison

Granite 4.0 Micro vs Grok-3 mini

Grok-3 mini is the stronger model overall, scoring 41.2 to 29.0 on the Noometry Index.

Last verified . 7 shared benchmarks.

Granite 4.0 Micro IBM

29.0

Rank #318 Confirmed

Grok-3 mini xAI

41.2

Rank #141 Confirmed

Summary

  • They share 7 benchmarks with published results for both. Granite 4.0 Micro scores higher in 1 category and Grok-3 mini in 4 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok-3 mini leads 46.4 to 9.9.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 2.8% for Granite 4.0 Micro and 77.8% for Grok-3 mini.
  • Granite 4.0 Micro has downloadable open weights; the other is API-only.

Side by side

Granite 4.0 Micro and Grok-3 mini specifications
Granite 4.0 MicroGrok-3 mini
ProviderIBMxAI
Noometry Index29.041.2
Released2025-10-022025-04-09
WeightsOpenProprietary
Context window131K—
Max output118K—
Input $ / M tokens$0.017—
Output $ / M tokens$0.11—
Results tracked835

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Granite 4.0 Micro: —, Grok-3 mini: 40.8 (#131)

Coding benchmarks
BenchmarkGranite 4.0 MicroGrok-3 mini
Aider Polyglot—49.3%
WeirdML—42.6%
LMArena Coding—1379

Reasoning Granite 4.0 Micro leads

Granite 4.0 Micro: 19.2 (#265), Grok-3 mini: 13.6 (#334)

Reasoning benchmarks
BenchmarkGranite 4.0 MicroGrok-3 mini
ARC-AGI-2—0.4%
Kagi LLM Benchmark—61.3%
ARC-AGI-1—16.5%
Chess Puzzles0%—
LMArena Hard Prompts—1375
Epoch Capabilities Index—140.35

Math Grok-3 mini leads

Granite 4.0 Micro: 12.0 (#307), Grok-3 mini: 42.1 (#85)

Math benchmarks
BenchmarkGranite 4.0 MicroGrok-3 mini
OTIS Mock AIME 2024-20252.8%77.8%
Omni-MATH20.9%31.8%
LMArena Math—1386
MATH Level 5—90.9%
FrontierMath (Feb 2025 set)—5.9%

Knowledge Grok-3 mini leads

Granite 4.0 Micro: 9.9 (#304), Grok-3 mini: 46.4 (#81)

Knowledge benchmarks
BenchmarkGranite 4.0 MicroGrok-3 mini
GPQA Diamond28.3%76.3%
MMLU-Pro39.5%79.9%
GPQA (HELM)30.7%67.5%
Confabulations—10.8%
LMArena Expert—1395

Multilingual Not comparable

Granite 4.0 Micro: —, Grok-3 mini: 48.1 (#145)

Multilingual benchmarks
BenchmarkGranite 4.0 MicroGrok-3 mini
LMArena Non-English—1352
LMArena Chinese—1387
LMArena French—1357
LMArena German—1349
LMArena Japanese—1342
LMArena Korean—1335
LMArena Russian—1353
LMArena Spanish—1381

Instruction Following Grok-3 mini leads

Granite 4.0 Micro: 69.9 (#169), Grok-3 mini: 78.5 (#9)

Instruction Following benchmarks
BenchmarkGranite 4.0 MicroGrok-3 mini
IFEval84.9%95.1%
LMArena Instruction Following—1357

Long Context Not comparable

Granite 4.0 Micro: —, Grok-3 mini: 41.0 (#147)

Long Context benchmarks
BenchmarkGranite 4.0 MicroGrok-3 mini
Fiction.LiveBench—66.7%
LMArena Longer Query—1372

Writing & Preference Grok-3 mini leads

Granite 4.0 Micro: 46.7 (#216), Grok-3 mini: 52.5 (#169)

Writing & Preference benchmarks
BenchmarkGranite 4.0 MicroGrok-3 mini
WildBench67%65.1%
LMArena Text—1370
LMArena Creative Writing—1342
Short-Story Creative Writing—73.5%
LMArena Multi-Turn—1355

Frequently asked questions

Is Granite 4.0 Micro better than Grok-3 mini?

Grok-3 mini is the stronger model overall, scoring 41.2 to 29.0 on the Noometry Index.

How many benchmarks do Granite 4.0 Micro and Grok-3 mini share?

7 benchmarks have published results for both models. Granite 4.0 Micro has 8 scored results on Noometry and Grok-3 mini has 35.

Related comparisons

Go deeper