Model comparison

Granite 4.0 Micro vs Kimi K2 Thinking Turbo

Kimi K2 Thinking Turbo is the stronger model overall, scoring 45.8 to 29.0 on the Noometry Index.

Last verified . 3 shared benchmarks.

Granite 4.0 Micro IBM

29.0

Rank #318 Confirmed

Kimi K2 Thinking Turbo Moonshot AI

45.8

Rank #70 Confirmed

Summary

  • They share 3 benchmarks with published results for both. Granite 4.0 Micro scores higher in 0 categories and Kimi K2 Thinking Turbo in 5 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Kimi K2 Thinking Turbo leads 50.9 to 9.9.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 2.8% for Granite 4.0 Micro and 83.1% for Kimi K2 Thinking Turbo.

Side by side

Granite 4.0 Micro and Kimi K2 Thinking Turbo specifications
Granite 4.0 MicroKimi K2 Thinking Turbo
ProviderIBMMoonshot AI
Noometry Index29.045.8
Released2025-10-022025-11-06
WeightsOpenOpen
Context window131K—
Max output118K—
Input $ / M tokens$0.017—
Output $ / M tokens$0.11—
Results tracked821

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Granite 4.0 Micro: —, Kimi K2 Thinking Turbo: 38.4 (#178)

Coding benchmarks
BenchmarkGranite 4.0 MicroKimi K2 Thinking Turbo
LMArena WebDev—1322
LMArena Coding—1454

Reasoning Kimi K2 Thinking Turbo leads

Granite 4.0 Micro: 19.2 (#265), Kimi K2 Thinking Turbo: 32.5 (#75)

Reasoning benchmarks
BenchmarkGranite 4.0 MicroKimi K2 Thinking Turbo
Chess Puzzles0%20%
LMArena Hard Prompts—1428

Math Kimi K2 Thinking Turbo leads

Granite 4.0 Micro: 12.0 (#307), Kimi K2 Thinking Turbo: 47.4 (#68)

Math benchmarks
BenchmarkGranite 4.0 MicroKimi K2 Thinking Turbo
OTIS Mock AIME 2024-20252.8%83.1%
Omni-MATH20.9%—
LMArena Math—1429

Knowledge Kimi K2 Thinking Turbo leads

Granite 4.0 Micro: 9.9 (#304), Kimi K2 Thinking Turbo: 50.9 (#69)

Knowledge benchmarks
BenchmarkGranite 4.0 MicroKimi K2 Thinking Turbo
GPQA Diamond28.3%84.2%
MMLU-Pro39.5%—
GPQA (HELM)30.7%—
LMArena Expert—1439

Multilingual Not comparable

Granite 4.0 Micro: —, Kimi K2 Thinking Turbo: 51.4 (#109)

Multilingual benchmarks
BenchmarkGranite 4.0 MicroKimi K2 Thinking Turbo
LMArena Non-English—1398
LMArena Chinese—1456
LMArena French—1425
LMArena German—1390
LMArena Japanese—1357
LMArena Korean—1330
LMArena Russian—1391
LMArena Spanish—1406

Instruction Following Kimi K2 Thinking Turbo leads

Granite 4.0 Micro: 69.9 (#169), Kimi K2 Thinking Turbo: 74.0 (#109)

Instruction Following benchmarks
BenchmarkGranite 4.0 MicroKimi K2 Thinking Turbo
IFEval84.9%—
LMArena Instruction Following—1403

Long Context Not comparable

Granite 4.0 Micro: —, Kimi K2 Thinking Turbo: 43.2 (#102)

Long Context benchmarks
BenchmarkGranite 4.0 MicroKimi K2 Thinking Turbo
LMArena Longer Query—1415

Writing & Preference Kimi K2 Thinking Turbo leads

Granite 4.0 Micro: 46.7 (#216), Kimi K2 Thinking Turbo: 60.0 (#104)

Writing & Preference benchmarks
BenchmarkGranite 4.0 MicroKimi K2 Thinking Turbo
LMArena Text—1415
LMArena Creative Writing—1374
WildBench67%—
LMArena Multi-Turn—1414

Frequently asked questions

Is Granite 4.0 Micro better than Kimi K2 Thinking Turbo?

Kimi K2 Thinking Turbo is the stronger model overall, scoring 45.8 to 29.0 on the Noometry Index.

How many benchmarks do Granite 4.0 Micro and Kimi K2 Thinking Turbo share?

3 benchmarks have published results for both models. Granite 4.0 Micro has 8 scored results on Noometry and Kimi K2 Thinking Turbo has 21.

Related comparisons

Go deeper