Model comparison

Grok 4.6 vs Kimi K2.5 Instant

Grok 4.6 is the stronger model overall, scoring 56.9 to 43.6 on the Noometry Index.

Last verified . 18 shared benchmarks.

Grok 4.6 xAI

56.9

Rank #21 Confirmed

Kimi K2.5 Instant Moonshot AI

43.6

Rank #89 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Grok 4.6 scores higher in 9 categories and Kimi K2.5 Instant in 0 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.6 leads 61.4 to 29.7.
  • Kimi K2.5 Instant has downloadable open weights; the other is API-only.

Side by side

Grok 4.6 and Kimi K2.5 Instant specifications
Grok 4.6Kimi K2.5 Instant
ProviderxAIMoonshot AI
Noometry Index56.943.6
Released2026-08-12—
WeightsProprietaryOpen
Context window500K—
Max output500K—
Input $ / M tokens$2—
Output $ / M tokens$6—
Results tracked4918

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.6 leads

Grok 4.6: 58.5 (#16), Kimi K2.5 Instant: 42.6 (#97)

Coding benchmarks
BenchmarkGrok 4.6Kimi K2.5 Instant
LMArena WebDev16171404
LMArena Coding14651484
DeepSWE67.5%—
FrontierCode48%—
CursorBench41.4%—
FrontierSWE25.3%—
SciCode56.5%—
WeirdML67.3%—
ALE-Bench1,508—

Agentic & Tool Use Not comparable

Grok 4.6: 39.4 (#27), Kimi K2.5 Instant: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.6Kimi K2.5 Instant
APEX-Agents65.3%—
GDP.pdf17.2%—
Vending-Bench 29,047—

Reasoning Grok 4.6 leads

Grok 4.6: 61.4 (#20), Kimi K2.5 Instant: 29.7 (#90)

Reasoning benchmarks
BenchmarkGrok 4.6Kimi K2.5 Instant
LMArena Hard Prompts14471443
ARC-AGI-267.1%—
SimpleBench75.9%—
NYT Connections (extended)80%—
ARC-AGI-187.5%—
CritPt19.7%—
Chess Puzzles40%—
EBR-Bench30.5%—
Mystery Game Puzzles34%—
DTBench97.3%—
LMCA48.5%—
Epoch Capabilities Index156.44—

Math Grok 4.6 leads

Grok 4.6: 67.0 (#24), Kimi K2.5 Instant: 39.4 (#105)

Math benchmarks
BenchmarkGrok 4.6Kimi K2.5 Instant
LMArena Math14231442
FrontierMath (Tiers 1-3)66%—
FrontierMath Tier 431.7%—
OTIS Mock AIME 2024-202599.2%—
ProofBench51%—

Knowledge Grok 4.6 leads

Grok 4.6: 63.3 (#20), Kimi K2.5 Instant: 40.2 (#123)

Knowledge benchmarks
BenchmarkGrok 4.6Kimi K2.5 Instant
LMArena Expert14671440
GPQA Diamond94%—
SimpleQA Verified49.3%—

Multimodal Grok 4.6 leads

Grok 4.6: 43.6 (#23), Kimi K2.5 Instant: 40.2 (#50)

Multimodal benchmarks
BenchmarkGrok 4.6Kimi K2.5 Instant
LMArena Vision12631254
Blueprint-Bench 233.2%—
Furniture Assembly40%—
LMArena Document1452—

Multilingual Too close to call

Grok 4.6: 53.0 (#74), Kimi K2.5 Instant: 52.0 (#94)

Multilingual benchmarks
BenchmarkGrok 4.6Kimi K2.5 Instant
LMArena Non-English14201406
LMArena Chinese14801449
LMArena French14611403
LMArena German14311413
LMArena Korean13971378
LMArena Russian14221404
LMArena Spanish14041447
LMArena Japanese1376—

Instruction Following Too close to call

Grok 4.6: 75.4 (#63), Kimi K2.5 Instant: 75.3 (#65)

Instruction Following benchmarks
BenchmarkGrok 4.6Kimi K2.5 Instant
LMArena Instruction Following14311430

Long Context Too close to call

Grok 4.6: 44.5 (#66), Kimi K2.5 Instant: 43.9 (#83)

Long Context benchmarks
BenchmarkGrok 4.6Kimi K2.5 Instant
LMArena Longer Query14541435

Writing & Preference Grok 4.6 leads

Grok 4.6: 62.3 (#80), Kimi K2.5 Instant: 60.6 (#95)

Writing & Preference benchmarks
BenchmarkGrok 4.6Kimi K2.5 Instant
LMArena Text14281420
LMArena Creative Writing14281381
LMArena Multi-Turn14251427

Frequently asked questions

Is Grok 4.6 better than Kimi K2.5 Instant?

Grok 4.6 is the stronger model overall, scoring 56.9 to 43.6 on the Noometry Index.

Is Grok 4.6 or Kimi K2.5 Instant better for coding?

Grok 4.6 scores higher on coding benchmarks: 58.5 versus 42.6 in the Noometry coding category.

How many benchmarks do Grok 4.6 and Kimi K2.5 Instant share?

18 benchmarks have published results for both models. Grok 4.6 has 49 scored results on Noometry and Kimi K2.5 Instant has 18.

Related comparisons

Go deeper