Model comparison

GPT-4.5 vs Kimi K2 (Jul 2025)

Kimi K2 (Jul 2025) is the stronger model overall, scoring 41.2 to 37.2 on the Noometry Index.

Last verified . 25 shared benchmarks.

GPT-4.5 OpenAI

37.2

Rank #208 Confirmed

Kimi K2 (Jul 2025) Moonshot AI

41.2

Rank #140 Confirmed

Summary

  • They share 25 benchmarks with published results for both. GPT-4.5 scores higher in 2 categories and Kimi K2 (Jul 2025) in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Kimi K2 (Jul 2025) leads 42.7 to 32.6.
  • The biggest single-benchmark swing is Aider Polyglot: 44.9% for GPT-4.5 and 59.1% for Kimi K2 (Jul 2025).
  • Kimi K2 (Jul 2025) has downloadable open weights; the other is API-only.

Side by side

GPT-4.5 and Kimi K2 (Jul 2025) specifications
GPT-4.5Kimi K2 (Jul 2025)
ProviderOpenAIMoonshot AI
Noometry Index37.241.2
Released2025-02-272025-07-12
WeightsProprietaryOpen
Context window—262K
Max output—262K
Input $ / M tokens—$0.57
Output $ / M tokens—$2.30
Results tracked4242

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

GPT-4.5: 42.2 (#109), Kimi K2 (Jul 2025): 42.4 (#102)

Coding benchmarks
BenchmarkGPT-4.5Kimi K2 (Jul 2025)
Aider Polyglot44.9%59.1%
WeirdML39.4%42.8%
LMArena Coding13961399
SWE-bench Verified (bash only)—63.4%
GSO—4.9%
LiveBench Coding75.2%—
ALE-Bench—597.5

Agentic & Tool Use Kimi K2 (Jul 2025) leads

GPT-4.5: 27.9 (#97), Kimi K2 (Jul 2025): 32.4 (#64)

Agentic & Tool Use benchmarks
BenchmarkGPT-4.5Kimi K2 (Jul 2025)
Terminal-Bench—35.7%
Berkeley Function Calling Leaderboard—59.1%
Cybench17.5%—
METR Time Horizons—59.2%

Reasoning Kimi K2 (Jul 2025) leads

GPT-4.5: 13.9 (#330), Kimi K2 (Jul 2025): 23.3 (#179)

Reasoning benchmarks
BenchmarkGPT-4.5Kimi K2 (Jul 2025)
SimpleBench34.5%26.3%
LMArena Hard Prompts14031384
Epoch Capabilities Index136.74146.01
ForecastBench61.760.2
ARC-AGI-20.8%—
Kagi LLM Benchmark—64.4%
ARC-AGI-110.3%—
EnigmaEval3.2%—
LiveBench Reasoning71.1%—
LiveBench Data Analysis64.3%—
LiveBench69%—

Math Kimi K2 (Jul 2025) leads

GPT-4.5: 32.6 (#211), Kimi K2 (Jul 2025): 42.7 (#83)

Math benchmarks
BenchmarkGPT-4.5Kimi K2 (Jul 2025)
LMArena Math14121397
OTIS Mock AIME 2024-202537.8%—
Omni-MATH—65.4%
LiveBench Math69.3%—
MATH Level 578.6%—
FrontierMath (Feb 2025 set)—21.4%
FrontierMath Tier 4 (v1)—0%

Knowledge Kimi K2 (Jul 2025) leads

GPT-4.5: 32.5 (#211), Kimi K2 (Jul 2025): 37.3 (#157)

Knowledge benchmarks
BenchmarkGPT-4.5Kimi K2 (Jul 2025)
Confabulations13.6%20.4%
LMArena Expert13941365
GPQA Diamond68.7%—
Humanity's Last Exam5.4%—
MMLU-Pro—81.9%
Vectara Hallucination Rate—17.9%
GPQA (HELM)—65.3%

Multimodal Not comparable

GPT-4.5: 37.6 (#71), Kimi K2 (Jul 2025): —

Multimodal benchmarks
BenchmarkGPT-4.5Kimi K2 (Jul 2025)
LMArena Vision1195—
VPCT45%—

Multilingual GPT-4.5 leads

GPT-4.5: 52.5 (#83), Kimi K2 (Jul 2025): 49.6 (#130)

Multilingual benchmarks
BenchmarkGPT-4.5Kimi K2 (Jul 2025)
LMArena Non-English14131372
LMArena Chinese14211415
LMArena French14181379
LMArena German14571387
LMArena Japanese14161349
LMArena Korean13921325
LMArena Russian14191385
LMArena Spanish—1386

Instruction Following GPT-4.5 leads

GPT-4.5: 72.6 (#134), Kimi K2 (Jul 2025): 71.1 (#156)

Instruction Following benchmarks
BenchmarkGPT-4.5Kimi K2 (Jul 2025)
LMArena Instruction Following14041348
LiveBench Instruction Following72.3%—
IFEval—85%

Long Context Too close to call

GPT-4.5: 40.4 (#155), Kimi K2 (Jul 2025): 41.2 (#145)

Long Context benchmarks
BenchmarkGPT-4.5Kimi K2 (Jul 2025)
Fiction.LiveBench63.9%66.7%
LMArena Longer Query14061353
CL-bench—17.6%

Writing & Preference Kimi K2 (Jul 2025) leads

GPT-4.5: 56.9 (#134), Kimi K2 (Jul 2025): 62.3 (#78)

Writing & Preference benchmarks
BenchmarkGPT-4.5Kimi K2 (Jul 2025)
LMArena Text14171380
LMArena Creative Writing13941350
Short-Story Creative Writing75.6%85.6%
EQ-Bench Creative Writing12581666
LMArena Multi-Turn14441371
WildBench—86.2%
LiveBench Language61.5%—

Frequently asked questions

Is GPT-4.5 better than Kimi K2 (Jul 2025)?

Kimi K2 (Jul 2025) is the stronger model overall, scoring 41.2 to 37.2 on the Noometry Index.

Is GPT-4.5 or Kimi K2 (Jul 2025) better for coding?

They score almost the same on coding (42.2 vs 42.4); test both on your own repository before choosing.

How many benchmarks do GPT-4.5 and Kimi K2 (Jul 2025) share?

25 benchmarks have published results for both models. GPT-4.5 has 42 scored results on Noometry and Kimi K2 (Jul 2025) has 42.

Related comparisons

Go deeper