Model comparison

Kimi K2.5 Instant vs o1

Kimi K2.5 Instant is the stronger model overall, scoring 43.6 to 40.9 on the Noometry Index.

Last verified . 17 shared benchmarks.

Kimi K2.5 Instant Moonshot AI

43.6

Rank #89 Confirmed

o1 OpenAI

40.9

Rank #143 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Kimi K2.5 Instant scores higher in 6 categories and o1 in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in long context, where o1 leads 50.3 to 43.9.
  • Kimi K2.5 Instant has downloadable open weights; the other is API-only.

Side by side

Kimi K2.5 Instant and o1 specifications
Kimi K2.5 Instanto1
ProviderMoonshot AIOpenAI
Noometry Index43.640.9
Released—2024-09-12
WeightsOpenProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$15
Output $ / M tokens—$60
Results tracked1852

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1 leads

Kimi K2.5 Instant: 42.6 (#97), o1: 46.1 (#70)

Coding benchmarks
BenchmarkKimi K2.5 Instanto1
LMArena Coding14841367
Aider Polyglot—61.7%
LMArena WebDev1404—
WeirdML—47.6%
LiveBench Coding—69.7%
CadEval—56%
HumanEval+—89%
MBPP+—80.2%

Agentic & Tool Use Not comparable

Kimi K2.5 Instant: —, o1: 24.6 (#117)

Agentic & Tool Use benchmarks
BenchmarkKimi K2.5 Instanto1
Cybench—10%
METR Time Horizons—51.1%

Reasoning Kimi K2.5 Instant leads

Kimi K2.5 Instant: 29.7 (#90), o1: 27.9 (#111)

Reasoning benchmarks
BenchmarkKimi K2.5 Instanto1
LMArena Hard Prompts14431371
SimpleBench—41.7%
ARC-AGI-1—30.7%
Chess Puzzles—15%
EnigmaEval—5.7%
LiveBench Reasoning—91.6%
DTBench—74.7%
LiveBench Data Analysis—65.5%
LMCA—22.3%
Epoch Capabilities Index—141.91
LiveBench—75.7%

Math Kimi K2.5 Instant leads

Kimi K2.5 Instant: 39.4 (#105), o1: 36.1 (#175)

Math benchmarks
BenchmarkKimi K2.5 Instanto1
LMArena Math14421388
FrontierMath (Tiers 1-3)—14.7%
OTIS Mock AIME 2024-2025—73.3%
LiveBench Math—80.3%
MATH Level 5—94.7%
FrontierMath (Feb 2025 set)—9.3%

Knowledge o1 leads

Kimi K2.5 Instant: 40.2 (#123), o1: 41.5 (#110)

Knowledge benchmarks
BenchmarkKimi K2.5 Instanto1
LMArena Expert14401361
GPQA Diamond—76.8%
Humanity's Last Exam—8%
SimpleQA Verified—41.1%
Confabulations—11.7%

Multimodal Kimi K2.5 Instant leads

Kimi K2.5 Instant: 40.2 (#50), o1: 34.2 (#93)

Multimodal benchmarks
BenchmarkKimi K2.5 Instanto1
LMArena Vision12541168
GeoBench—80%
VPCT—37%
SpatialViz-Bench—41.4%

Multilingual Kimi K2.5 Instant leads

Kimi K2.5 Instant: 52.0 (#94), o1: 48.6 (#142)

Multilingual benchmarks
BenchmarkKimi K2.5 Instanto1
LMArena Non-English14061358
LMArena Chinese14491394
LMArena French14031344
LMArena German14131337
LMArena Korean13781396
LMArena Russian14041356
LMArena Spanish14471345
LMArena Japanese—1346

Instruction Following Too close to call

Kimi K2.5 Instant: 75.3 (#65), o1: 74.8 (#86)

Instruction Following benchmarks
BenchmarkKimi K2.5 Instanto1
LMArena Instruction Following14301367
LiveBench Instruction Following—81.5%

Long Context o1 leads

Kimi K2.5 Instant: 43.9 (#83), o1: 50.3 (#9)

Long Context benchmarks
BenchmarkKimi K2.5 Instanto1
LMArena Longer Query14351378
Fiction.LiveBench—83.3%

Writing & Preference Kimi K2.5 Instant leads

Kimi K2.5 Instant: 60.6 (#95), o1: 55.6 (#144)

Writing & Preference benchmarks
BenchmarkKimi K2.5 Instanto1
LMArena Text14201366
LMArena Creative Writing13811348
LMArena Multi-Turn14271369
Short-Story Creative Writing—70.2%
LiveBench Language—65.4%

Frequently asked questions

Is Kimi K2.5 Instant better than o1?

Kimi K2.5 Instant is the stronger model overall, scoring 43.6 to 40.9 on the Noometry Index.

Is Kimi K2.5 Instant or o1 better for coding?

o1 scores higher on coding benchmarks: 46.1 versus 42.6 in the Noometry coding category.

How many benchmarks do Kimi K2.5 Instant and o1 share?

17 benchmarks have published results for both models. Kimi K2.5 Instant has 18 scored results on Noometry and o1 has 52.

Related comparisons

Go deeper