Model comparison

GPT-5.4 nano vs Kimi K2.7 Code

Kimi K2.7 Code is the stronger model overall, scoring 43.3 to 41.9 on the Noometry Index. GPT-5.4 nano costs 3.7× less per token, which makes it the better buy when Kimi K2.7 Code's lead doesn't matter for your workload.

Last verified . 11 shared benchmarks.

GPT-5.4 nano OpenAI

41.9

Rank #125 Confirmed

Kimi K2.7 Code Moonshot AI

43.3

Rank #94 Confirmed

Summary

  • They share 11 benchmarks with published results for both. GPT-5.4 nano scores higher in 1 category and Kimi K2.7 Code in 3 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Kimi K2.7 Code leads 39.0 to 23.7.
  • The biggest single-benchmark swing is SimpleQA Verified: 11.7% for GPT-5.4 nano and 36.5% for Kimi K2.7 Code.
  • GPT-5.4 nano is cheaper at $0.20 / $1.25 per million input/output tokens, against $0.95 / $4 for Kimi K2.7 Code.
  • GPT-5.4 nano accepts more context: 400K tokens versus 262K.
  • Kimi K2.7 Code has downloadable open weights; the other is API-only.

Side by side

GPT-5.4 nano and Kimi K2.7 Code specifications
GPT-5.4 nanoKimi K2.7 Code
ProviderOpenAIMoonshot AI
Noometry Index41.943.3
Released2026-03-172026-06-12
WeightsProprietaryOpen
Context window400K262K
Max output128K262K
Input $ / M tokens$0.20$0.95
Output $ / M tokens$1.25$4
Results tracked4019

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

GPT-5.4 nano: 43.6 (#84), Kimi K2.7 Code: 42.9 (#95)

Coding benchmarks
BenchmarkGPT-5.4 nanoKimi K2.7 Code
SciCode46.9%47.5%
WeirdML49.2%54.1%
ALE-Bench1,005886.23
DeepSWE—30.5%
FrontierCode—30.1%
LMArena WebDev—1473
LMArena Coding1405—

Agentic & Tool Use Not comparable

GPT-5.4 nano: —, Kimi K2.7 Code: 24.0 (#122)

Agentic & Tool Use benchmarks
BenchmarkGPT-5.4 nanoKimi K2.7 Code
APEX-Agents—37.6%
GBAEval—0.9%
Vending-Bench 2—5,083

Reasoning Kimi K2.7 Code leads

GPT-5.4 nano: 23.7 (#173), Kimi K2.7 Code: 39.0 (#61)

Reasoning benchmarks
BenchmarkGPT-5.4 nanoKimi K2.7 Code
CritPt9.3%10%
Chess Puzzles30%21%
Epoch Capabilities Index145.81149.97
ARC-AGI-25.7%—
SimpleBench—57.9%
Kagi LLM Benchmark39.7%—
ARC-AGI-151.5%—
LMArena Hard Prompts1381—
Mystery Game Puzzles9%—
DTBench80.3%—
LMCA36.9%—
Surface Evolver Bench—48.8%
ForecastBench57.3—

Math Kimi K2.7 Code leads

GPT-5.4 nano: 40.9 (#88), Kimi K2.7 Code: 52.9 (#48)

Math benchmarks
BenchmarkGPT-5.4 nanoKimi K2.7 Code
FrontierMath (Tiers 1-3)44.9%54%
FrontierMath Tier 412.2%12.2%
OTIS Mock AIME 2024-202587.8%95.6%
ProofBench5%—
LMArena Math1406—
FrontierMath (Feb 2025 set)25.9%—
FrontierMath Tier 4 (v1)6.3%—

Knowledge Kimi K2.7 Code leads

GPT-5.4 nano: 41.9 (#103), Kimi K2.7 Code: 53.5 (#57)

Knowledge benchmarks
BenchmarkGPT-5.4 nanoKimi K2.7 Code
GPQA Diamond78.5%87.9%
SimpleQA Verified11.7%36.5%
Vectara Hallucination Rate3.1%—
LMArena Expert1396—

Multimodal Not comparable

GPT-5.4 nano: 36.7 (#78), Kimi K2.7 Code: —

Multimodal benchmarks
BenchmarkGPT-5.4 nanoKimi K2.7 Code
LMArena Vision1196—

Multilingual Not comparable

GPT-5.4 nano: 48.6 (#140), Kimi K2.7 Code: —

Multilingual benchmarks
BenchmarkGPT-5.4 nanoKimi K2.7 Code
LMArena Non-English1359—
LMArena Chinese1392—
LMArena French1396—
LMArena German1367—
LMArena Japanese1343—
LMArena Korean1320—
LMArena Russian1363—
LMArena Spanish1371—

Instruction Following Not comparable

GPT-5.4 nano: 71.9 (#144), Kimi K2.7 Code: —

Instruction Following benchmarks
BenchmarkGPT-5.4 nanoKimi K2.7 Code
LMArena Instruction Following1362—

Long Context Not comparable

GPT-5.4 nano: 41.6 (#137), Kimi K2.7 Code: —

Long Context benchmarks
BenchmarkGPT-5.4 nanoKimi K2.7 Code
LMArena Longer Query1366—

Writing & Preference Not comparable

GPT-5.4 nano: 55.7 (#142), Kimi K2.7 Code: —

Writing & Preference benchmarks
BenchmarkGPT-5.4 nanoKimi K2.7 Code
LMArena Text1372—
LMArena Creative Writing1314—
LMArena Multi-Turn1382—

Frequently asked questions

Is GPT-5.4 nano better than Kimi K2.7 Code?

Kimi K2.7 Code is the stronger model overall, scoring 43.3 to 41.9 on the Noometry Index. GPT-5.4 nano costs 3.7× less per token, which makes it the better buy when Kimi K2.7 Code's lead doesn't matter for your workload.

Which is cheaper, GPT-5.4 nano or Kimi K2.7 Code?

GPT-5.4 nano is cheaper. It lists at $0.20 per million input tokens and $1.25 per million output tokens; Kimi K2.7 Code lists at $0.95 and $4.

Is GPT-5.4 nano or Kimi K2.7 Code better for coding?

They score almost the same on coding (43.6 vs 42.9); test both on your own repository before choosing.

Which has the bigger context window?

GPT-5.4 nano does, with 400K tokens against 262K.

How many benchmarks do GPT-5.4 nano and Kimi K2.7 Code share?

11 benchmarks have published results for both models. GPT-5.4 nano has 40 scored results on Noometry and Kimi K2.7 Code has 19.

Related comparisons

Go deeper