Model comparison

GPT-5 Mini vs Kimi K2.5

Kimi K2.5 is the stronger model overall, scoring 48.1 to 41.8 on the Noometry Index.

Last verified . 42 shared benchmarks.

GPT-5 Mini OpenAI

41.8

Rank #128 Confirmed

Kimi K2.5 Moonshot AI

48.1

Rank #57 Confirmed

Summary

  • They share 42 benchmarks with published results for both. GPT-5 Mini scores higher in 1 category and Kimi K2.5 in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Kimi K2.5 leads 52.1 to 41.9.
  • The biggest single-benchmark swing is SWE-bench Multilingual: 39.7% for GPT-5 Mini and 67.3% for Kimi K2.5.
  • GPT-5 Mini is cheaper at $0.25 / $2 per million input/output tokens, against $0.45 / $2.25 for Kimi K2.5.
  • GPT-5 Mini accepts more context: 400K tokens versus 262K.
  • Kimi K2.5 has downloadable open weights; the other is API-only.

Side by side

GPT-5 Mini and Kimi K2.5 specifications
GPT-5 MiniKimi K2.5
ProviderOpenAIMoonshot AI
Noometry Index41.848.1
Released2025-08-072026-01-27
WeightsProprietaryOpen
Context window400K262K
Max output128K262K
Input $ / M tokens$0.25$0.45
Output $ / M tokens$2$2.25
Results tracked6051

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2.5 leads

GPT-5 Mini: 40.1 (#146), Kimi K2.5: 48.8 (#53)

Coding benchmarks
BenchmarkGPT-5 MiniKimi K2.5
SWE-bench Verified64.7%73.8%
SWE-bench Verified (bash only)59.8%70.8%
SWE-bench Multilingual39.7%67.3%
SciCode39.2%49%
WeirdML52.7%45.6%
LMArena Coding14061474
ALE-Bench799.77821.65
LMArena WebDev—1437
AlgoTune1.38—

Agentic & Tool Use Kimi K2.5 leads

GPT-5 Mini: 31.1 (#70), Kimi K2.5: 34.2 (#48)

Agentic & Tool Use benchmarks
BenchmarkGPT-5 MiniKimi K2.5
Terminal-Bench34.8%43.2%
Vending-Bench 2-31.181,198
Berkeley Function Calling Leaderboard55.5%—
OSWorld—63.3%

Reasoning Kimi K2.5 leads

GPT-5 Mini: 23.9 (#168), Kimi K2.5: 31.2 (#80)

Reasoning benchmarks
BenchmarkGPT-5 MiniKimi K2.5
ARC-AGI-24.4%11.8%
Kagi LLM Benchmark70.3%78.5%
ARC-AGI-154.3%65.3%
CritPt0%3.1%
Chess Puzzles30%12%
EnigmaEval8.2%3.4%
LMArena Hard Prompts13801453
Epoch Capabilities Index145.52148.03
SimpleBench—46.8%
NYT Connections (extended)—69.9%
Thematic Generalization—69.4%
Mystery Game Puzzles10%—
DTBench80.5%—
LMCA34.2%—
ForecastBench61—

Math Kimi K2.5 leads

GPT-5 Mini: 46.7 (#69), Kimi K2.5: 51.8 (#53)

Knowledge Kimi K2.5 leads

GPT-5 Mini: 45.6 (#86), Kimi K2.5: 53.6 (#56)

Knowledge benchmarks
BenchmarkGPT-5 MiniKimi K2.5
GPQA Diamond75%87.6%
Humanity's Last Exam19.4%24.4%
SimpleQA Verified21.6%34.3%
Vectara Hallucination Rate12.9%14.2%
LMArena Expert13791466
MMLU-Pro83.5%—
Confabulations13.3%—
GPQA (HELM)75.6%—

Multimodal Kimi K2.5 leads

GPT-5 Mini: 35.6 (#85), Kimi K2.5: 41.1 (#39)

Multimodal benchmarks
BenchmarkGPT-5 MiniKimi K2.5
LMArena Vision12021269
VPCT40.2%—
LMArena Document—1430

Multilingual Kimi K2.5 leads

GPT-5 Mini: 48.9 (#137), Kimi K2.5: 53.9 (#53)

Multilingual benchmarks
BenchmarkGPT-5 MiniKimi K2.5
LMArena Non-English13631433
LMArena Chinese13851495
LMArena French13861454
LMArena German13661441
LMArena Japanese13411421
LMArena Korean13081410
LMArena Russian13621435
LMArena Spanish13551450

Instruction Following Too close to call

GPT-5 Mini: 76.2 (#46), Kimi K2.5: 75.3 (#64)

Instruction Following benchmarks
BenchmarkGPT-5 MiniKimi K2.5
LMArena Instruction Following13571431
IFEval92.7%—

Long Context Kimi K2.5 leads

GPT-5 Mini: 41.9 (#132), Kimi K2.5: 52.1 (#7)

Long Context benchmarks
BenchmarkGPT-5 MiniKimi K2.5
Fiction.LiveBench69.4%86.1%
LMArena Longer Query13551445
CL-bench—19.3%
CL-bench Life—13.2%

Writing & Preference Kimi K2.5 leads

GPT-5 Mini: 55.2 (#148), Kimi K2.5: 65.1 (#53)

Writing & Preference benchmarks
BenchmarkGPT-5 MiniKimi K2.5
LMArena Text13731445
LMArena Creative Writing13251423
EQ-Bench Creative Writing13131579
LMArena Multi-Turn13631444
Short-Story Creative Writing83.1%—
WildBench85.5%—

Frequently asked questions

Is GPT-5 Mini better than Kimi K2.5?

Kimi K2.5 is the stronger model overall, scoring 48.1 to 41.8 on the Noometry Index.

Which is cheaper, GPT-5 Mini or Kimi K2.5?

GPT-5 Mini is cheaper. It lists at $0.25 per million input tokens and $2 per million output tokens; Kimi K2.5 lists at $0.45 and $2.25.

Is GPT-5 Mini or Kimi K2.5 better for coding?

Kimi K2.5 scores higher on coding benchmarks: 48.8 versus 40.1 in the Noometry coding category.

Which has the bigger context window?

GPT-5 Mini does, with 400K tokens against 262K.

How many benchmarks do GPT-5 Mini and Kimi K2.5 share?

42 benchmarks have published results for both models. GPT-5 Mini has 60 scored results on Noometry and Kimi K2.5 has 51.

Related comparisons

Go deeper