Model comparison

GPT-4.1 nano vs Kimi K2 (Jul 2025)

Kimi K2 (Jul 2025) is the stronger model overall, scoring 41.2 to 27.9 on the Noometry Index. GPT-4.1 nano costs 5.7× less per token, which makes it the better buy when Kimi K2 (Jul 2025)'s lead doesn't matter for your workload.

Last verified . 27 shared benchmarks.

GPT-4.1 nano OpenAI

27.9

Rank #327 Confirmed

Kimi K2 (Jul 2025) Moonshot AI

41.2

Rank #140 Confirmed

Summary

  • They share 27 benchmarks with published results for both. GPT-4.1 nano scores higher in 0 categories and Kimi K2 (Jul 2025) in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Kimi K2 (Jul 2025) leads 62.3 to 40.5.
  • The biggest single-benchmark swing is Aider Polyglot: 8.9% for GPT-4.1 nano and 59.1% for Kimi K2 (Jul 2025).
  • GPT-4.1 nano is cheaper at $0.10 / $0.40 per million input/output tokens, against $0.57 / $2.30 for Kimi K2 (Jul 2025).
  • GPT-4.1 nano accepts more context: 1.05M tokens versus 262K.
  • Kimi K2 (Jul 2025) has downloadable open weights; the other is API-only.

Side by side

GPT-4.1 nano and Kimi K2 (Jul 2025) specifications
GPT-4.1 nanoKimi K2 (Jul 2025)
ProviderOpenAIMoonshot AI
Noometry Index27.941.2
Released2025-04-142025-07-12
WeightsProprietaryOpen
Context window1.05M262K
Max output33K262K
Input $ / M tokens$0.10$0.57
Output $ / M tokens$0.40$2.30
Results tracked3842

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2 (Jul 2025) leads

GPT-4.1 nano: 24.1 (#330), Kimi K2 (Jul 2025): 42.4 (#102)

Coding benchmarks
BenchmarkGPT-4.1 nanoKimi K2 (Jul 2025)
Aider Polyglot8.9%59.1%
WeirdML19%42.8%
LMArena Coding13061399
SWE-bench Verified (bash only)—63.4%
SciCode25.9%—
GSO—4.9%
ALE-Bench—597.5

Agentic & Tool Use Kimi K2 (Jul 2025) leads

GPT-4.1 nano: 26.5 (#104), Kimi K2 (Jul 2025): 32.4 (#64)

Agentic & Tool Use benchmarks
BenchmarkGPT-4.1 nanoKimi K2 (Jul 2025)
Berkeley Function Calling Leaderboard33%59.1%
Terminal-Bench—35.7%
METR Time Horizons—59.2%

Reasoning Kimi K2 (Jul 2025) leads

GPT-4.1 nano: 8.5 (#349), Kimi K2 (Jul 2025): 23.3 (#179)

Reasoning benchmarks
BenchmarkGPT-4.1 nanoKimi K2 (Jul 2025)
Kagi LLM Benchmark33.3%64.4%
LMArena Hard Prompts12861384
Epoch Capabilities Index129.62146.01
ARC-AGI-20%—
SimpleBench—26.3%
ARC-AGI-10%—
CritPt0%—
DTBench52.5%—
LMCA5.5%—
ForecastBench—60.2

Math Kimi K2 (Jul 2025) leads

GPT-4.1 nano: 26.9 (#252), Kimi K2 (Jul 2025): 42.7 (#83)

Math benchmarks
BenchmarkGPT-4.1 nanoKimi K2 (Jul 2025)
Omni-MATH36.7%65.4%
LMArena Math12741397
FrontierMath (Feb 2025 set)1%21.4%
OTIS Mock AIME 2024-202528.9%—
MATH Level 570%—
FrontierMath Tier 4 (v1)—0%

Knowledge Kimi K2 (Jul 2025) leads

GPT-4.1 nano: 21.8 (#273), Kimi K2 (Jul 2025): 37.3 (#157)

Knowledge benchmarks
BenchmarkGPT-4.1 nanoKimi K2 (Jul 2025)
MMLU-Pro55%81.9%
GPQA (HELM)50.7%65.3%
LMArena Expert12721365
GPQA Diamond48.9%—
SimpleQA Verified6%—
Confabulations—20.4%
Vectara Hallucination Rate—17.9%

Multimodal Not comparable

GPT-4.1 nano: 29.2 (#113), Kimi K2 (Jul 2025): —

Multimodal benchmarks
BenchmarkGPT-4.1 nanoKimi K2 (Jul 2025)
LMArena Vision1063—

Multilingual Kimi K2 (Jul 2025) leads

GPT-4.1 nano: 41.6 (#205), Kimi K2 (Jul 2025): 49.6 (#130)

Multilingual benchmarks
BenchmarkGPT-4.1 nanoKimi K2 (Jul 2025)
LMArena Non-English12601372
LMArena Chinese12701415
LMArena German12881387
LMArena Japanese11981349
LMArena Russian12611385
LMArena French—1379
LMArena Korean—1325
LMArena Spanish—1386

Instruction Following Kimi K2 (Jul 2025) leads

GPT-4.1 nano: 67.8 (#193), Kimi K2 (Jul 2025): 71.1 (#156)

Instruction Following benchmarks
BenchmarkGPT-4.1 nanoKimi K2 (Jul 2025)
IFEval84.3%85%
LMArena Instruction Following12671348

Long Context Kimi K2 (Jul 2025) leads

GPT-4.1 nano: 23.7 (#296), Kimi K2 (Jul 2025): 41.2 (#145)

Long Context benchmarks
BenchmarkGPT-4.1 nanoKimi K2 (Jul 2025)
Fiction.LiveBench25%66.7%
LMArena Longer Query12831353
CL-bench—17.6%

Writing & Preference Kimi K2 (Jul 2025) leads

GPT-4.1 nano: 40.5 (#243), Kimi K2 (Jul 2025): 62.3 (#78)

Writing & Preference benchmarks
BenchmarkGPT-4.1 nanoKimi K2 (Jul 2025)
LMArena Text12851380
LMArena Creative Writing12601350
EQ-Bench Creative Writing9461666
WildBench81.2%86.2%
LMArena Multi-Turn12771371
Short-Story Creative Writing—85.6%

Frequently asked questions

Is GPT-4.1 nano better than Kimi K2 (Jul 2025)?

Kimi K2 (Jul 2025) is the stronger model overall, scoring 41.2 to 27.9 on the Noometry Index. GPT-4.1 nano costs 5.7× less per token, which makes it the better buy when Kimi K2 (Jul 2025)'s lead doesn't matter for your workload.

Which is cheaper, GPT-4.1 nano or Kimi K2 (Jul 2025)?

GPT-4.1 nano is cheaper. It lists at $0.10 per million input tokens and $0.40 per million output tokens; Kimi K2 (Jul 2025) lists at $0.57 and $2.30.

Is GPT-4.1 nano or Kimi K2 (Jul 2025) better for coding?

Kimi K2 (Jul 2025) scores higher on coding benchmarks: 42.4 versus 24.1 in the Noometry coding category.

Which has the bigger context window?

GPT-4.1 nano does, with 1.05M tokens against 262K.

How many benchmarks do GPT-4.1 nano and Kimi K2 (Jul 2025) share?

27 benchmarks have published results for both models. GPT-4.1 nano has 38 scored results on Noometry and Kimi K2 (Jul 2025) has 42.

Related comparisons

Go deeper