Model comparison

Grok 4.20 (Non-Reasoning) vs Kimi K3

Kimi K3 is the stronger model overall, scoring 59.5 to 48.6 on the Noometry Index. Grok 4.20 (Non-Reasoning) costs 3.8× less per token, which makes it the better buy when Kimi K3's lead doesn't matter for your workload.

Last verified . 38 shared benchmarks.

Grok 4.20 (Non-Reasoning) xAI

48.6

Rank #54 Confirmed

Kimi K3 Moonshot AI

59.5

Rank #15 Confirmed

Summary

  • They share 38 benchmarks with published results for both. Grok 4.20 (Non-Reasoning) scores higher in 0 categories and Kimi K3 in 10 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Kimi K3 leads 74.2 to 48.2.
  • The biggest single-benchmark swing is ProofBench: 14% for Grok 4.20 (Non-Reasoning) and 87% for Kimi K3.
  • Grok 4.20 (Non-Reasoning) is cheaper at $1.25 / $2.50 per million input/output tokens, against $3 / $15 for Kimi K3.
  • Kimi K3 accepts more context: 1.05M tokens versus 1M.
  • Kimi K3 has downloadable open weights; the other is API-only.

Side by side

Grok 4.20 (Non-Reasoning) and Kimi K3 specifications
Grok 4.20 (Non-Reasoning)Kimi K3
ProviderxAIMoonshot AI
Noometry Index48.659.5
Released2026-02-172026-07-16
WeightsProprietaryOpen
Context window1M1.05M
Max output30K1.05M
Input $ / M tokens$1.25$3
Output $ / M tokens$2.50$15
Results tracked4653

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K3 leads

Grok 4.20 (Non-Reasoning): 42.1 (#112), Kimi K3: 61.0 (#10)

Coding benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Kimi K3
LMArena WebDev13751654
WeirdML52.3%82.6%
LMArena Coding14591508
ALE-Bench1,1501,524
DeepSWE—68.5%
FrontierCode—44.2%
FrontierSWE—25.9%
SciCode—59.5%

Agentic & Tool Use Kimi K3 leads

Grok 4.20 (Non-Reasoning): 34.4 (#46), Kimi K3: 41.8 (#20)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Kimi K3
τ²-bench Banking18%37.1%
Vending-Bench 24,6635,165
Terminal-Bench57.3%—
APEX-Agents—50.6%
PostTrainBench—32%
GBAEval—48.3%
GDP.pdf—19%
LMArena Search1189—

Reasoning Kimi K3 leads

Grok 4.20 (Non-Reasoning): 52.3 (#32), Kimi K3: 63.0 (#17)

Reasoning benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Kimi K3
ARC-AGI-265.1%60.4%
NYT Connections (extended)85.4%93.6%
ARC-AGI-189.5%94.5%
Chess Puzzles24%39%
LMArena Hard Prompts14511496
DTBench90.1%91.2%
LMCA38.7%52.7%
Epoch Capabilities Index151.98157.45
ForecastBench61.461.1
SimpleBench—60.7%
Kagi LLM Benchmark75%—
CritPt—23.4%
Thematic Generalization63.8%—
Mystery Game Puzzles—26%
Surface Evolver Bench—95%

Math Kimi K3 leads

Grok 4.20 (Non-Reasoning): 48.2 (#65), Kimi K3: 74.2 (#16)

Math benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Kimi K3
FrontierMath (Tiers 1-3)44.9%72.2%
FrontierMath Tier 417.1%39%
OTIS Mock AIME 2024-202592.2%97.2%
ProofBench14%87%
LMArena Math14551491
MathArena Final-Answer Competitions—87.8%

Knowledge Kimi K3 leads

Grok 4.20 (Non-Reasoning): 52.8 (#60), Kimi K3: 63.2 (#21)

Knowledge benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Kimi K3
GPQA Diamond89.3%93.1%
SimpleQA Verified30.2%50.6%
LMArena Expert14391521

Multimodal Kimi K3 leads

Grok 4.20 (Non-Reasoning): 33.3 (#98), Kimi K3: 37.8 (#70)

Multimodal benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Kimi K3
Blueprint-Bench 20%29.5%
LMArena Vision1263—
Furniture Assembly—34.2%
LMArena Document1416—

Multilingual Kimi K3 leads

Grok 4.20 (Non-Reasoning): 54.5 (#40), Kimi K3: 56.3 (#21)

Multilingual benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Kimi K3
LMArena Non-English14411466
LMArena Chinese14811529
LMArena French14761491
LMArena German14651488
LMArena Japanese14491487
LMArena Korean14171458
LMArena Russian14581482
LMArena Spanish14431472

Instruction Following Kimi K3 leads

Grok 4.20 (Non-Reasoning): 74.8 (#83), Kimi K3: 77.7 (#14)

Instruction Following benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Kimi K3
LMArena Instruction Following14201483

Long Context Too close to call

Grok 4.20 (Non-Reasoning): 45.5 (#34), Kimi K3: 45.8 (#29)

Long Context benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Kimi K3
LMArena Longer Query14371494
CL-bench22.2%—
CL-bench Life11.9%—

Writing & Preference Kimi K3 leads

Grok 4.20 (Non-Reasoning): 65.7 (#44), Kimi K3: 76.6 (#4)

Writing & Preference benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Kimi K3
LMArena Text14511476
LMArena Creative Writing14381454
EQ-Bench Creative Writing15742082
LMArena Multi-Turn14561488
EQ-Bench 4—1339

Frequently asked questions

Is Grok 4.20 (Non-Reasoning) better than Kimi K3?

Kimi K3 is the stronger model overall, scoring 59.5 to 48.6 on the Noometry Index. Grok 4.20 (Non-Reasoning) costs 3.8× less per token, which makes it the better buy when Kimi K3's lead doesn't matter for your workload.

Which is cheaper, Grok 4.20 (Non-Reasoning) or Kimi K3?

Grok 4.20 (Non-Reasoning) is cheaper. It lists at $1.25 per million input tokens and $2.50 per million output tokens; Kimi K3 lists at $3 and $15.

Is Grok 4.20 (Non-Reasoning) or Kimi K3 better for coding?

Kimi K3 scores higher on coding benchmarks: 61.0 versus 42.1 in the Noometry coding category.

Which has the bigger context window?

Kimi K3 does, with 1.05M tokens against 1M.

How many benchmarks do Grok 4.20 (Non-Reasoning) and Kimi K3 share?

38 benchmarks have published results for both models. Grok 4.20 (Non-Reasoning) has 46 scored results on Noometry and Kimi K3 has 53.

Related comparisons

Go deeper