Model comparison

GPT-4.1 vs Kimi K2.6

Kimi K2.6 is the stronger model overall, scoring 47.7 to 35.9 on the Noometry Index.

Last verified . 33 shared benchmarks.

GPT-4.1 OpenAI

35.9

Rank #219 Confirmed

Kimi K2.6 Moonshot AI

47.7

Rank #60 Confirmed

Summary

  • They share 33 benchmarks with published results for both. GPT-4.1 scores higher in 2 categories and Kimi K2.6 in 8 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where Kimi K2.6 leads 57.0 to 22.3.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 38.3% for GPT-4.1 and 96.1% for Kimi K2.6.
  • Kimi K2.6 is cheaper at $0.95 / $4 per million input/output tokens, against $2 / $8 for GPT-4.1.
  • GPT-4.1 accepts more context: 1.05M tokens versus 262K.
  • Kimi K2.6 has downloadable open weights; the other is API-only.

Side by side

GPT-4.1 and Kimi K2.6 specifications
GPT-4.1Kimi K2.6
ProviderOpenAIMoonshot AI
Noometry Index35.947.7
Released2025-04-142026-04-20
WeightsProprietaryOpen
Context window1.05M262K
Max output33K262K
Input $ / M tokens$2$0.95
Output $ / M tokens$8$4
Results tracked5251

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2.6 leads

GPT-4.1: 34.4 (#238), Kimi K2.6: 50.7 (#43)

Coding benchmarks
BenchmarkGPT-4.1Kimi K2.6
SWE-bench Verified48.5%76.7%
WeirdML39%55.9%
LMArena Coding13911488
ALE-Bench558.11,093
SWE-bench Verified (bash only)39.6%—
Aider Polyglot52.4%—
LMArena WebDev—1509
SciCode—53.5%
CadEval42%—

Agentic & Tool Use GPT-4.1 leads

GPT-4.1: 34.7 (#43), Kimi K2.6: 21.9 (#137)

Agentic & Tool Use benchmarks
BenchmarkGPT-4.1Kimi K2.6
Berkeley Function Calling Leaderboard54%—
OSWorld 2.0—4.6%
ExploitBench—18.4%
GBAEval—0.9%
GDP.pdf—12%
Vending-Bench 2—6,205

Reasoning Kimi K2.6 leads

GPT-4.1: 11.7 (#339), Kimi K2.6: 40.5 (#55)

Reasoning benchmarks
BenchmarkGPT-4.1Kimi K2.6
Chess Puzzles6%26%
LMArena Hard Prompts13841470
DTBench68.3%90.9%
LMCA25.6%37.3%
Epoch Capabilities Index136.78151.05
ARC-AGI-20.4%—
SimpleBench27%—
Kagi LLM Benchmark52.3%—
NYT Connections (extended)—87.2%
ARC-AGI-15.5%—
CritPt—8%
EnigmaEval2.2%—
EBR-Bench—2.4%
Mystery Game Puzzles—18%
ForecastBench61.5—

Math Kimi K2.6 leads

GPT-4.1: 22.3 (#280), Kimi K2.6: 57.0 (#41)

Knowledge Kimi K2.6 leads

GPT-4.1: 37.1 (#160), Kimi K2.6: 54.0 (#54)

Knowledge benchmarks
BenchmarkGPT-4.1Kimi K2.6
GPQA Diamond66.9%90.8%
SimpleQA Verified31.1%34.9%
Vectara Hallucination Rate5.6%10.8%
LMArena Expert13641491
Humanity's Last Exam5.4%—
MMLU-Pro81.1%—
GPQA (HELM)65.9%—

Multimodal GPT-4.1 leads

GPT-4.1: 38.2 (#67), Kimi K2.6: 31.6 (#103)

Multimodal benchmarks
BenchmarkGPT-4.1Kimi K2.6
LMArena Vision12111283
GeoBench72%—
Blueprint-Bench 2—3.9%
Furniture Assembly—21.7%
LMArena Document—1451

Multilingual Kimi K2.6 leads

GPT-4.1: 49.4 (#133), Kimi K2.6: 54.9 (#37)

Multilingual benchmarks
BenchmarkGPT-4.1Kimi K2.6
LMArena Non-English13701446
LMArena Chinese13821521
LMArena French13821471
LMArena German13811450
LMArena Japanese13191443
LMArena Korean13391427
LMArena Russian13771446
LMArena Spanish13761464

Instruction Following Kimi K2.6 leads

GPT-4.1: 71.3 (#153), Kimi K2.6: 76.3 (#43)

Instruction Following benchmarks
BenchmarkGPT-4.1Kimi K2.6
LMArena Instruction Following13671451
IFEval83.8%—

Long Context Kimi K2.6 leads

GPT-4.1: 40.0 (#163), Kimi K2.6: 44.9 (#52)

Long Context benchmarks
BenchmarkGPT-4.1Kimi K2.6
LMArena Longer Query13851468
Fiction.LiveBench63.9%—

Writing & Preference Kimi K2.6 leads

GPT-4.1: 57.6 (#125), Kimi K2.6: 68.5 (#26)

Writing & Preference benchmarks
BenchmarkGPT-4.1Kimi K2.6
LMArena Text13831455
LMArena Creative Writing13631434
EQ-Bench Creative Writing14201725
LMArena Multi-Turn13981453
WildBench85.4%—
EQ-Bench 4—1202

Frequently asked questions

Is GPT-4.1 better than Kimi K2.6?

Kimi K2.6 is the stronger model overall, scoring 47.7 to 35.9 on the Noometry Index.

Which is cheaper, GPT-4.1 or Kimi K2.6?

Kimi K2.6 is cheaper. It lists at $0.95 per million input tokens and $4 per million output tokens; GPT-4.1 lists at $2 and $8.

Is GPT-4.1 or Kimi K2.6 better for coding?

Kimi K2.6 scores higher on coding benchmarks: 50.7 versus 34.4 in the Noometry coding category.

Which has the bigger context window?

GPT-4.1 does, with 1.05M tokens against 262K.

How many benchmarks do GPT-4.1 and Kimi K2.6 share?

33 benchmarks have published results for both models. GPT-4.1 has 52 scored results on Noometry and Kimi K2.6 has 51.

Related comparisons

Go deeper