Model comparison

GPT-5.5 Instant vs Kimi K2.7 Code

GPT-5.5 Instant and Kimi K2.7 Code score almost the same on the Noometry Index (42.7 vs 43.3), so choose on price, context window or the category you care about most.

Last verified . 8 shared benchmarks.

GPT-5.5 Instant OpenAI

42.7

Rank #110 Confirmed

Kimi K2.7 Code Moonshot AI

43.3

Rank #94 Confirmed

Summary

  • They share 8 benchmarks with published results for both. GPT-5.5 Instant scores higher in 1 category and Kimi K2.7 Code in 3 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in math, where Kimi K2.7 Code leads 52.9 to 26.5.
  • The biggest single-benchmark swing is FrontierMath (Tiers 1-3): 26.3% for GPT-5.5 Instant and 54% for Kimi K2.7 Code.
  • Kimi K2.7 Code has downloadable open weights; the other is API-only.

Side by side

GPT-5.5 Instant and Kimi K2.7 Code specifications
GPT-5.5 InstantKimi K2.7 Code
ProviderOpenAIMoonshot AI
Noometry Index42.743.3
Released2026-05-052026-06-12
WeightsProprietaryOpen
Context window—262K
Max output—262K
Input $ / M tokens—$0.95
Output $ / M tokens—$4
Results tracked2719

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.5 Instant leads

GPT-5.5 Instant: 44.3 (#74), Kimi K2.7 Code: 42.9 (#95)

Coding benchmarks
BenchmarkGPT-5.5 InstantKimi K2.7 Code
SciCode48.6%47.5%
DeepSWE—30.5%
FrontierCode—30.1%
LMArena WebDev—1473
WeirdML—54.1%
LMArena Coding1433—
ALE-Bench—886.23

Agentic & Tool Use Not comparable

GPT-5.5 Instant: —, Kimi K2.7 Code: 24.0 (#122)

Agentic & Tool Use benchmarks
BenchmarkGPT-5.5 InstantKimi K2.7 Code
APEX-Agents—37.6%
GBAEval—0.9%
Vending-Bench 2—5,083

Reasoning Kimi K2.7 Code leads

GPT-5.5 Instant: 24.9 (#155), Kimi K2.7 Code: 39.0 (#61)

Reasoning benchmarks
BenchmarkGPT-5.5 InstantKimi K2.7 Code
CritPt0%10%
Chess Puzzles12%21%
Epoch Capabilities Index142.52149.97
SimpleBench—57.9%
LMArena Hard Prompts1426—
Surface Evolver Bench—48.8%

Math Kimi K2.7 Code leads

GPT-5.5 Instant: 26.5 (#259), Kimi K2.7 Code: 52.9 (#48)

Math benchmarks
BenchmarkGPT-5.5 InstantKimi K2.7 Code
FrontierMath (Tiers 1-3)26.3%54%
FrontierMath Tier 42.4%12.2%
OTIS Mock AIME 2024-202568.1%95.6%
LMArena Math1420—

Knowledge Kimi K2.7 Code leads

GPT-5.5 Instant: 48.9 (#74), Kimi K2.7 Code: 53.5 (#57)

Knowledge benchmarks
BenchmarkGPT-5.5 InstantKimi K2.7 Code
GPQA Diamond82.5%87.9%
SimpleQA Verified—36.5%
LMArena Expert1409—

Multimodal Not comparable

GPT-5.5 Instant: 40.0 (#52), Kimi K2.7 Code: —

Multimodal benchmarks
BenchmarkGPT-5.5 InstantKimi K2.7 Code
LMArena Vision1250—
LMArena Document1403—

Multilingual Not comparable

GPT-5.5 Instant: 52.8 (#80), Kimi K2.7 Code: —

Multilingual benchmarks
BenchmarkGPT-5.5 InstantKimi K2.7 Code
LMArena Non-English1417—
LMArena Chinese1456—
LMArena French1428—
LMArena German1411—
LMArena Japanese1408—
LMArena Korean1392—
LMArena Russian1431—
LMArena Spanish1429—

Instruction Following Not comparable

GPT-5.5 Instant: 74.2 (#100), Kimi K2.7 Code: —

Instruction Following benchmarks
BenchmarkGPT-5.5 InstantKimi K2.7 Code
LMArena Instruction Following1406—

Long Context Not comparable

GPT-5.5 Instant: 43.4 (#96), Kimi K2.7 Code: —

Long Context benchmarks
BenchmarkGPT-5.5 InstantKimi K2.7 Code
LMArena Longer Query1422—

Writing & Preference Not comparable

GPT-5.5 Instant: 61.8 (#85), Kimi K2.7 Code: —

Writing & Preference benchmarks
BenchmarkGPT-5.5 InstantKimi K2.7 Code
LMArena Text1419—
LMArena Creative Writing1419—
LMArena Multi-Turn1433—

Frequently asked questions

Is GPT-5.5 Instant better than Kimi K2.7 Code?

GPT-5.5 Instant and Kimi K2.7 Code score almost the same on the Noometry Index (42.7 vs 43.3), so choose on price, context window or the category you care about most.

Is GPT-5.5 Instant or Kimi K2.7 Code better for coding?

GPT-5.5 Instant scores higher on coding benchmarks: 44.3 versus 42.9 in the Noometry coding category.

How many benchmarks do GPT-5.5 Instant and Kimi K2.7 Code share?

8 benchmarks have published results for both models. GPT-5.5 Instant has 27 scored results on Noometry and Kimi K2.7 Code has 19.

Related comparisons

Go deeper