Model comparison

GPT-5.4 nano vs Kimi K2 (Jul 2025)

GPT-5.4 nano and Kimi K2 (Jul 2025) score almost the same on the Noometry Index (41.9 vs 41.2), so choose on price, context window or the category you care about most.

Last verified . 25 shared benchmarks.

GPT-5.4 nano OpenAI

41.9

Rank #125 Confirmed

Kimi K2 (Jul 2025) Moonshot AI

41.2

Rank #140 Confirmed

Summary

  • They share 25 benchmarks with published results for both. GPT-5.4 nano scores higher in 5 categories and Kimi K2 (Jul 2025) in 3 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Kimi K2 (Jul 2025) leads 62.3 to 55.7.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 39.7% for GPT-5.4 nano and 64.4% for Kimi K2 (Jul 2025).
  • GPT-5.4 nano is cheaper at $0.20 / $1.25 per million input/output tokens, against $0.57 / $2.30 for Kimi K2 (Jul 2025).
  • GPT-5.4 nano accepts more context: 400K tokens versus 262K.
  • Kimi K2 (Jul 2025) has downloadable open weights; the other is API-only.

Side by side

GPT-5.4 nano and Kimi K2 (Jul 2025) specifications
GPT-5.4 nanoKimi K2 (Jul 2025)
ProviderOpenAIMoonshot AI
Noometry Index41.941.2
Released2026-03-172025-07-12
WeightsProprietaryOpen
Context window400K262K
Max output128K262K
Input $ / M tokens$0.20$0.57
Output $ / M tokens$1.25$2.30
Results tracked4042

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.4 nano leads

GPT-5.4 nano: 43.6 (#84), Kimi K2 (Jul 2025): 42.4 (#102)

Coding benchmarks
BenchmarkGPT-5.4 nanoKimi K2 (Jul 2025)
WeirdML49.2%42.8%
LMArena Coding14051399
ALE-Bench1,005597.5
SWE-bench Verified (bash only)—63.4%
Aider Polyglot—59.1%
SciCode46.9%—
GSO—4.9%

Agentic & Tool Use Not comparable

GPT-5.4 nano: —, Kimi K2 (Jul 2025): 32.4 (#64)

Agentic & Tool Use benchmarks
BenchmarkGPT-5.4 nanoKimi K2 (Jul 2025)
Terminal-Bench—35.7%
Berkeley Function Calling Leaderboard—59.1%
METR Time Horizons—59.2%

Reasoning Too close to call

GPT-5.4 nano: 23.7 (#173), Kimi K2 (Jul 2025): 23.3 (#179)

Reasoning benchmarks
BenchmarkGPT-5.4 nanoKimi K2 (Jul 2025)
Kagi LLM Benchmark39.7%64.4%
LMArena Hard Prompts13811384
Epoch Capabilities Index145.81146.01
ForecastBench57.360.2
ARC-AGI-25.7%—
SimpleBench—26.3%
ARC-AGI-151.5%—
CritPt9.3%—
Chess Puzzles30%—
Mystery Game Puzzles9%—
DTBench80.3%—
LMCA36.9%—

Math Kimi K2 (Jul 2025) leads

GPT-5.4 nano: 40.9 (#88), Kimi K2 (Jul 2025): 42.7 (#83)

Math benchmarks
BenchmarkGPT-5.4 nanoKimi K2 (Jul 2025)
LMArena Math14061397
FrontierMath (Feb 2025 set)25.9%21.4%
FrontierMath Tier 4 (v1)6.3%0%
FrontierMath (Tiers 1-3)44.9%—
FrontierMath Tier 412.2%—
OTIS Mock AIME 2024-202587.8%—
ProofBench5%—
Omni-MATH—65.4%

Knowledge GPT-5.4 nano leads

GPT-5.4 nano: 41.9 (#103), Kimi K2 (Jul 2025): 37.3 (#157)

Knowledge benchmarks
BenchmarkGPT-5.4 nanoKimi K2 (Jul 2025)
Vectara Hallucination Rate3.1%17.9%
LMArena Expert13961365
GPQA Diamond78.5%—
SimpleQA Verified11.7%—
MMLU-Pro—81.9%
Confabulations—20.4%
GPQA (HELM)—65.3%

Multimodal Not comparable

GPT-5.4 nano: 36.7 (#78), Kimi K2 (Jul 2025): —

Multimodal benchmarks
BenchmarkGPT-5.4 nanoKimi K2 (Jul 2025)
LMArena Vision1196—

Multilingual Too close to call

GPT-5.4 nano: 48.6 (#140), Kimi K2 (Jul 2025): 49.6 (#130)

Multilingual benchmarks
BenchmarkGPT-5.4 nanoKimi K2 (Jul 2025)
LMArena Non-English13591372
LMArena Chinese13921415
LMArena French13961379
LMArena German13671387
LMArena Japanese13431349
LMArena Korean13201325
LMArena Russian13631385
LMArena Spanish13711386

Instruction Following Too close to call

GPT-5.4 nano: 71.9 (#144), Kimi K2 (Jul 2025): 71.1 (#156)

Instruction Following benchmarks
BenchmarkGPT-5.4 nanoKimi K2 (Jul 2025)
LMArena Instruction Following13621348
IFEval—85%

Long Context Too close to call

GPT-5.4 nano: 41.6 (#137), Kimi K2 (Jul 2025): 41.2 (#145)

Long Context benchmarks
BenchmarkGPT-5.4 nanoKimi K2 (Jul 2025)
LMArena Longer Query13661353
Fiction.LiveBench—66.7%
CL-bench—17.6%

Writing & Preference Kimi K2 (Jul 2025) leads

GPT-5.4 nano: 55.7 (#142), Kimi K2 (Jul 2025): 62.3 (#78)

Writing & Preference benchmarks
BenchmarkGPT-5.4 nanoKimi K2 (Jul 2025)
LMArena Text13721380
LMArena Creative Writing13141350
LMArena Multi-Turn13821371
Short-Story Creative Writing—85.6%
EQ-Bench Creative Writing—1666
WildBench—86.2%

Frequently asked questions

Is GPT-5.4 nano better than Kimi K2 (Jul 2025)?

GPT-5.4 nano and Kimi K2 (Jul 2025) score almost the same on the Noometry Index (41.9 vs 41.2), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-5.4 nano or Kimi K2 (Jul 2025)?

GPT-5.4 nano is cheaper. It lists at $0.20 per million input tokens and $1.25 per million output tokens; Kimi K2 (Jul 2025) lists at $0.57 and $2.30.

Is GPT-5.4 nano or Kimi K2 (Jul 2025) better for coding?

GPT-5.4 nano scores higher on coding benchmarks: 43.6 versus 42.4 in the Noometry coding category.

Which has the bigger context window?

GPT-5.4 nano does, with 400K tokens against 262K.

How many benchmarks do GPT-5.4 nano and Kimi K2 (Jul 2025) share?

25 benchmarks have published results for both models. GPT-5.4 nano has 40 scored results on Noometry and Kimi K2 (Jul 2025) has 42.

Related comparisons

Go deeper