Model comparison

Grok 4.1 vs Kimi K2.5

Kimi K2.5 is the stronger model overall, scoring 48.1 to 41.5 on the Noometry Index.

Last verified . 18 shared benchmarks.

Grok 4.1 xAI

41.5

Rank #134 Confirmed

Kimi K2.5 Moonshot AI

48.1

Rank #57 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Grok 4.1 scores higher in 0 categories and Kimi K2.5 in 9 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Kimi K2.5 leads 48.8 to 33.7.
  • Kimi K2.5 has downloadable open weights; the other is API-only.

Side by side

Grok 4.1 and Kimi K2.5 specifications
Grok 4.1Kimi K2.5
ProviderxAIMoonshot AI
Noometry Index41.548.1
Released2025-11-172026-01-27
WeightsProprietaryOpen
Context window—262K
Max output—262K
Input $ / M tokens—$0.45
Output $ / M tokens—$2.25
Results tracked1951

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2.5 leads

Grok 4.1: 33.7 (#253), Kimi K2.5: 48.8 (#53)

Coding benchmarks
BenchmarkGrok 4.1Kimi K2.5
LMArena WebDev12141437
LMArena Coding14451474
SWE-bench Verified—73.8%
SWE-bench Verified (bash only)—70.8%
SWE-bench Multilingual—67.3%
SciCode—49%
WeirdML—45.6%
ALE-Bench—821.65

Agentic & Tool Use Too close to call

Grok 4.1: 34.1 (#49), Kimi K2.5: 34.2 (#48)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.1Kimi K2.5
Terminal-Bench—43.2%
Cybench39%—
OSWorld—63.3%
Vending-Bench 2—1,198

Reasoning Kimi K2.5 leads

Grok 4.1: 29.5 (#91), Kimi K2.5: 31.2 (#80)

Reasoning benchmarks
BenchmarkGrok 4.1Kimi K2.5
LMArena Hard Prompts14351453
ARC-AGI-2—11.8%
SimpleBench—46.8%
Kagi LLM Benchmark—78.5%
NYT Connections (extended)—69.9%
ARC-AGI-1—65.3%
CritPt—3.1%
Chess Puzzles—12%
EnigmaEval—3.4%
Thematic Generalization—69.4%
Epoch Capabilities Index—148.03

Math Kimi K2.5 leads

Grok 4.1: 38.9 (#120), Kimi K2.5: 51.8 (#53)

Knowledge Kimi K2.5 leads

Grok 4.1: 39.5 (#133), Kimi K2.5: 53.6 (#56)

Knowledge benchmarks
BenchmarkGrok 4.1Kimi K2.5
LMArena Expert14171466
GPQA Diamond—87.6%
Humanity's Last Exam—24.4%
SimpleQA Verified—34.3%
Vectara Hallucination Rate—14.2%

Multimodal Not comparable

Grok 4.1: —, Kimi K2.5: 41.1 (#39)

Multimodal benchmarks
BenchmarkGrok 4.1Kimi K2.5
LMArena Vision—1269
LMArena Document—1430

Multilingual Too close to call

Grok 4.1: 53.4 (#68), Kimi K2.5: 53.9 (#53)

Multilingual benchmarks
BenchmarkGrok 4.1Kimi K2.5
LMArena Non-English14251433
LMArena Chinese14651495
LMArena French14481454
LMArena German14461441
LMArena Japanese13971421
LMArena Korean14071410
LMArena Russian14341435
LMArena Spanish14381450

Instruction Following Kimi K2.5 leads

Grok 4.1: 73.8 (#111), Kimi K2.5: 75.3 (#64)

Instruction Following benchmarks
BenchmarkGrok 4.1Kimi K2.5
LMArena Instruction Following14001431

Long Context Kimi K2.5 leads

Grok 4.1: 43.2 (#100), Kimi K2.5: 52.1 (#7)

Long Context benchmarks
BenchmarkGrok 4.1Kimi K2.5
LMArena Longer Query14161445
Fiction.LiveBench—86.1%
CL-bench—19.3%
CL-bench Life—13.2%

Writing & Preference Kimi K2.5 leads

Grok 4.1: 62.4 (#75), Kimi K2.5: 65.1 (#53)

Writing & Preference benchmarks
BenchmarkGrok 4.1Kimi K2.5
LMArena Text14371445
LMArena Creative Writing14111423
LMArena Multi-Turn14371444
EQ-Bench Creative Writing—1579

Frequently asked questions

Is Grok 4.1 better than Kimi K2.5?

Kimi K2.5 is the stronger model overall, scoring 48.1 to 41.5 on the Noometry Index.

Is Grok 4.1 or Kimi K2.5 better for coding?

Kimi K2.5 scores higher on coding benchmarks: 48.8 versus 33.7 in the Noometry coding category.

How many benchmarks do Grok 4.1 and Kimi K2.5 share?

18 benchmarks have published results for both models. Grok 4.1 has 19 scored results on Noometry and Kimi K2.5 has 51.

Related comparisons

Go deeper