Model comparison

Grok 4.1 vs Kimi K2.6

Kimi K2.6 is the stronger model overall, scoring 47.7 to 41.5 on the Noometry Index.

Last verified . 18 shared benchmarks.

Grok 4.1 xAI

41.5

Rank #134 Confirmed

Kimi K2.6 Moonshot AI

47.7

Rank #60 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Grok 4.1 scores higher in 1 category and Kimi K2.6 in 8 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Kimi K2.6 leads 57.0 to 38.9.
  • Kimi K2.6 has downloadable open weights; the other is API-only.

Side by side

Grok 4.1 and Kimi K2.6 specifications
Grok 4.1Kimi K2.6
ProviderxAIMoonshot AI
Noometry Index41.547.7
Released2025-11-172026-04-20
WeightsProprietaryOpen
Context window—262K
Max output—262K
Input $ / M tokens—$0.95
Output $ / M tokens—$4
Results tracked1951

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2.6 leads

Grok 4.1: 33.7 (#253), Kimi K2.6: 50.7 (#43)

Coding benchmarks
BenchmarkGrok 4.1Kimi K2.6
LMArena WebDev12141509
LMArena Coding14451488
SWE-bench Verified—76.7%
SciCode—53.5%
WeirdML—55.9%
ALE-Bench—1,093

Agentic & Tool Use Grok 4.1 leads

Grok 4.1: 34.1 (#49), Kimi K2.6: 21.9 (#137)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.1Kimi K2.6
OSWorld 2.0—4.6%
Cybench39%—
ExploitBench—18.4%
GBAEval—0.9%
GDP.pdf—12%
Vending-Bench 2—6,205

Reasoning Kimi K2.6 leads

Grok 4.1: 29.5 (#91), Kimi K2.6: 40.5 (#55)

Reasoning benchmarks
BenchmarkGrok 4.1Kimi K2.6
LMArena Hard Prompts14351470
NYT Connections (extended)—87.2%
CritPt—8%
Chess Puzzles—26%
EBR-Bench—2.4%
Mystery Game Puzzles—18%
DTBench—90.9%
LMCA—37.3%
Epoch Capabilities Index—151.05

Math Kimi K2.6 leads

Grok 4.1: 38.9 (#120), Kimi K2.6: 57.0 (#41)

Knowledge Kimi K2.6 leads

Grok 4.1: 39.5 (#133), Kimi K2.6: 54.0 (#54)

Knowledge benchmarks
BenchmarkGrok 4.1Kimi K2.6
LMArena Expert14171491
GPQA Diamond—90.8%
SimpleQA Verified—34.9%
Vectara Hallucination Rate—10.8%

Multimodal Not comparable

Grok 4.1: —, Kimi K2.6: 31.6 (#103)

Multimodal benchmarks
BenchmarkGrok 4.1Kimi K2.6
LMArena Vision—1283
Blueprint-Bench 2—3.9%
Furniture Assembly—21.7%
LMArena Document—1451

Multilingual Kimi K2.6 leads

Grok 4.1: 53.4 (#68), Kimi K2.6: 54.9 (#37)

Multilingual benchmarks
BenchmarkGrok 4.1Kimi K2.6
LMArena Non-English14251446
LMArena Chinese14651521
LMArena French14481471
LMArena German14461450
LMArena Japanese13971443
LMArena Korean14071427
LMArena Russian14341446
LMArena Spanish14381464

Instruction Following Kimi K2.6 leads

Grok 4.1: 73.8 (#111), Kimi K2.6: 76.3 (#43)

Instruction Following benchmarks
BenchmarkGrok 4.1Kimi K2.6
LMArena Instruction Following14001451

Long Context Kimi K2.6 leads

Grok 4.1: 43.2 (#100), Kimi K2.6: 44.9 (#52)

Long Context benchmarks
BenchmarkGrok 4.1Kimi K2.6
LMArena Longer Query14161468

Writing & Preference Kimi K2.6 leads

Grok 4.1: 62.4 (#75), Kimi K2.6: 68.5 (#26)

Writing & Preference benchmarks
BenchmarkGrok 4.1Kimi K2.6
LMArena Text14371455
LMArena Creative Writing14111434
LMArena Multi-Turn14371453
EQ-Bench Creative Writing—1725
EQ-Bench 4—1202

Frequently asked questions

Is Grok 4.1 better than Kimi K2.6?

Kimi K2.6 is the stronger model overall, scoring 47.7 to 41.5 on the Noometry Index.

Is Grok 4.1 or Kimi K2.6 better for coding?

Kimi K2.6 scores higher on coding benchmarks: 50.7 versus 33.7 in the Noometry coding category.

How many benchmarks do Grok 4.1 and Kimi K2.6 share?

18 benchmarks have published results for both models. Grok 4.1 has 19 scored results on Noometry and Kimi K2.6 has 51.

Related comparisons

Go deeper