Model comparison

GPT-5 Mini vs Kimi K2 (Jul 2025)

GPT-5 Mini and Kimi K2 (Jul 2025) score almost the same on the Noometry Index (41.8 vs 41.2), so choose on price, context window or the category you care about most.

Last verified . 37 shared benchmarks.

GPT-5 Mini OpenAI

41.8

Rank #128 Confirmed

Kimi K2 (Jul 2025) Moonshot AI

41.2

Rank #140 Confirmed

Summary

  • They share 37 benchmarks with published results for both. GPT-5 Mini scores higher in 5 categories and Kimi K2 (Jul 2025) in 4 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where GPT-5 Mini leads 45.6 to 37.3.
  • The biggest single-benchmark swing is GPQA (HELM): 75.6% for GPT-5 Mini and 65.3% for Kimi K2 (Jul 2025).
  • GPT-5 Mini is cheaper at $0.25 / $2 per million input/output tokens, against $0.57 / $2.30 for Kimi K2 (Jul 2025).
  • GPT-5 Mini accepts more context: 400K tokens versus 262K.
  • Kimi K2 (Jul 2025) has downloadable open weights; the other is API-only.

Side by side

GPT-5 Mini and Kimi K2 (Jul 2025) specifications
GPT-5 MiniKimi K2 (Jul 2025)
ProviderOpenAIMoonshot AI
Noometry Index41.841.2
Released2025-08-072025-07-12
WeightsProprietaryOpen
Context window400K262K
Max output128K262K
Input $ / M tokens$0.25$0.57
Output $ / M tokens$2$2.30
Results tracked6042

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2 (Jul 2025) leads

GPT-5 Mini: 40.1 (#146), Kimi K2 (Jul 2025): 42.4 (#102)

Coding benchmarks
BenchmarkGPT-5 MiniKimi K2 (Jul 2025)
SWE-bench Verified (bash only)59.8%63.4%
WeirdML52.7%42.8%
LMArena Coding14061399
ALE-Bench799.77597.5
SWE-bench Verified64.7%—
Aider Polyglot—59.1%
SWE-bench Multilingual39.7%—
SciCode39.2%—
GSO—4.9%
AlgoTune1.38—

Agentic & Tool Use Kimi K2 (Jul 2025) leads

GPT-5 Mini: 31.1 (#70), Kimi K2 (Jul 2025): 32.4 (#64)

Agentic & Tool Use benchmarks
BenchmarkGPT-5 MiniKimi K2 (Jul 2025)
Terminal-Bench34.8%35.7%
Berkeley Function Calling Leaderboard55.5%59.1%
METR Time Horizons—59.2%
Vending-Bench 2-31.18—

Reasoning Too close to call

GPT-5 Mini: 23.9 (#168), Kimi K2 (Jul 2025): 23.3 (#179)

Reasoning benchmarks
BenchmarkGPT-5 MiniKimi K2 (Jul 2025)
Kagi LLM Benchmark70.3%64.4%
LMArena Hard Prompts13801384
Epoch Capabilities Index145.52146.01
ForecastBench6160.2
ARC-AGI-24.4%—
SimpleBench—26.3%
ARC-AGI-154.3%—
CritPt0%—
Chess Puzzles30%—
EnigmaEval8.2%—
Mystery Game Puzzles10%—
DTBench80.5%—
LMCA34.2%—

Math GPT-5 Mini leads

GPT-5 Mini: 46.7 (#69), Kimi K2 (Jul 2025): 42.7 (#83)

Math benchmarks
BenchmarkGPT-5 MiniKimi K2 (Jul 2025)
Omni-MATH72.2%65.4%
LMArena Math13781397
FrontierMath (Feb 2025 set)27.2%21.4%
FrontierMath Tier 4 (v1)6.3%0%
FrontierMath (Tiers 1-3)46.7%—
FrontierMath Tier 412.2%—
OTIS Mock AIME 2024-202586.7%—
ProofBench9%—
MATH Level 597.8%—

Knowledge GPT-5 Mini leads

GPT-5 Mini: 45.6 (#86), Kimi K2 (Jul 2025): 37.3 (#157)

Knowledge benchmarks
BenchmarkGPT-5 MiniKimi K2 (Jul 2025)
MMLU-Pro83.5%81.9%
Confabulations13.3%20.4%
Vectara Hallucination Rate12.9%17.9%
GPQA (HELM)75.6%65.3%
LMArena Expert13791365
GPQA Diamond75%—
Humanity's Last Exam19.4%—
SimpleQA Verified21.6%—

Multimodal Not comparable

GPT-5 Mini: 35.6 (#85), Kimi K2 (Jul 2025): —

Multimodal benchmarks
BenchmarkGPT-5 MiniKimi K2 (Jul 2025)
LMArena Vision1202—
VPCT40.2%—

Multilingual Too close to call

GPT-5 Mini: 48.9 (#137), Kimi K2 (Jul 2025): 49.6 (#130)

Multilingual benchmarks
BenchmarkGPT-5 MiniKimi K2 (Jul 2025)
LMArena Non-English13631372
LMArena Chinese13851415
LMArena French13861379
LMArena German13661387
LMArena Japanese13411349
LMArena Korean13081325
LMArena Russian13621385
LMArena Spanish13551386

Instruction Following GPT-5 Mini leads

GPT-5 Mini: 76.2 (#46), Kimi K2 (Jul 2025): 71.1 (#156)

Instruction Following benchmarks
BenchmarkGPT-5 MiniKimi K2 (Jul 2025)
IFEval92.7%85%
LMArena Instruction Following13571348

Long Context Too close to call

GPT-5 Mini: 41.9 (#132), Kimi K2 (Jul 2025): 41.2 (#145)

Long Context benchmarks
BenchmarkGPT-5 MiniKimi K2 (Jul 2025)
Fiction.LiveBench69.4%66.7%
LMArena Longer Query13551353
CL-bench—17.6%

Writing & Preference Kimi K2 (Jul 2025) leads

GPT-5 Mini: 55.2 (#148), Kimi K2 (Jul 2025): 62.3 (#78)

Writing & Preference benchmarks
BenchmarkGPT-5 MiniKimi K2 (Jul 2025)
LMArena Text13731380
LMArena Creative Writing13251350
Short-Story Creative Writing83.1%85.6%
EQ-Bench Creative Writing13131666
WildBench85.5%86.2%
LMArena Multi-Turn13631371

Frequently asked questions

Is GPT-5 Mini better than Kimi K2 (Jul 2025)?

GPT-5 Mini and Kimi K2 (Jul 2025) score almost the same on the Noometry Index (41.8 vs 41.2), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-5 Mini or Kimi K2 (Jul 2025)?

GPT-5 Mini is cheaper. It lists at $0.25 per million input tokens and $2 per million output tokens; Kimi K2 (Jul 2025) lists at $0.57 and $2.30.

Is GPT-5 Mini or Kimi K2 (Jul 2025) better for coding?

Kimi K2 (Jul 2025) scores higher on coding benchmarks: 42.4 versus 40.1 in the Noometry coding category.

Which has the bigger context window?

GPT-5 Mini does, with 400K tokens against 262K.

How many benchmarks do GPT-5 Mini and Kimi K2 (Jul 2025) share?

37 benchmarks have published results for both models. GPT-5 Mini has 60 scored results on Noometry and Kimi K2 (Jul 2025) has 42.

Related comparisons

Go deeper