Model comparison

Kimi K2 Thinking Turbo vs o1

Kimi K2 Thinking Turbo is the stronger model overall, scoring 45.8 to 40.9 on the Noometry Index.

Last verified . 20 shared benchmarks.

Kimi K2 Thinking Turbo Moonshot AI

45.8

Rank #70 Confirmed

o1 OpenAI

40.9

Rank #143 Confirmed

Summary

  • They share 20 benchmarks with published results for both. Kimi K2 Thinking Turbo scores higher in 5 categories and o1 in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Kimi K2 Thinking Turbo leads 47.4 to 36.1.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 83.1% for Kimi K2 Thinking Turbo and 73.3% for o1.
  • Kimi K2 Thinking Turbo has downloadable open weights; the other is API-only.

Side by side

Kimi K2 Thinking Turbo and o1 specifications
Kimi K2 Thinking Turboo1
ProviderMoonshot AIOpenAI
Noometry Index45.840.9
Released2025-11-062024-09-12
WeightsOpenProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$15
Output $ / M tokens—$60
Results tracked2152

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1 leads

Kimi K2 Thinking Turbo: 38.4 (#178), o1: 46.1 (#70)

Coding benchmarks
BenchmarkKimi K2 Thinking Turboo1
LMArena Coding14541367
Aider Polyglot—61.7%
LMArena WebDev1322—
WeirdML—47.6%
LiveBench Coding—69.7%
CadEval—56%
HumanEval+—89%
MBPP+—80.2%

Agentic & Tool Use Not comparable

Kimi K2 Thinking Turbo: —, o1: 24.6 (#117)

Agentic & Tool Use benchmarks
BenchmarkKimi K2 Thinking Turboo1
Cybench—10%
METR Time Horizons—51.1%

Reasoning Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 32.5 (#75), o1: 27.9 (#111)

Reasoning benchmarks
BenchmarkKimi K2 Thinking Turboo1
Chess Puzzles20%15%
LMArena Hard Prompts14281371
SimpleBench—41.7%
ARC-AGI-1—30.7%
EnigmaEval—5.7%
LiveBench Reasoning—91.6%
DTBench—74.7%
LiveBench Data Analysis—65.5%
LMCA—22.3%
Epoch Capabilities Index—141.91
LiveBench—75.7%

Math Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 47.4 (#68), o1: 36.1 (#175)

Math benchmarks
BenchmarkKimi K2 Thinking Turboo1
OTIS Mock AIME 2024-202583.1%73.3%
LMArena Math14291388
FrontierMath (Tiers 1-3)—14.7%
LiveBench Math—80.3%
MATH Level 5—94.7%
FrontierMath (Feb 2025 set)—9.3%

Knowledge Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 50.9 (#69), o1: 41.5 (#110)

Knowledge benchmarks
BenchmarkKimi K2 Thinking Turboo1
GPQA Diamond84.2%76.8%
LMArena Expert14391361
Humanity's Last Exam—8%
SimpleQA Verified—41.1%
Confabulations—11.7%

Multimodal Not comparable

Kimi K2 Thinking Turbo: —, o1: 34.2 (#93)

Multimodal benchmarks
BenchmarkKimi K2 Thinking Turboo1
LMArena Vision—1168
GeoBench—80%
VPCT—37%
SpatialViz-Bench—41.4%

Multilingual Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 51.4 (#109), o1: 48.6 (#142)

Multilingual benchmarks
BenchmarkKimi K2 Thinking Turboo1
LMArena Non-English13981358
LMArena Chinese14561394
LMArena French14251344
LMArena German13901337
LMArena Japanese13571346
LMArena Korean13301396
LMArena Russian13911356
LMArena Spanish14061345

Instruction Following Too close to call

Kimi K2 Thinking Turbo: 74.0 (#109), o1: 74.8 (#86)

Instruction Following benchmarks
BenchmarkKimi K2 Thinking Turboo1
LMArena Instruction Following14031367
LiveBench Instruction Following—81.5%

Long Context o1 leads

Kimi K2 Thinking Turbo: 43.2 (#102), o1: 50.3 (#9)

Long Context benchmarks
BenchmarkKimi K2 Thinking Turboo1
LMArena Longer Query14151378
Fiction.LiveBench—83.3%

Writing & Preference Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 60.0 (#104), o1: 55.6 (#144)

Writing & Preference benchmarks
BenchmarkKimi K2 Thinking Turboo1
LMArena Text14151366
LMArena Creative Writing13741348
LMArena Multi-Turn14141369
Short-Story Creative Writing—70.2%
LiveBench Language—65.4%

Frequently asked questions

Is Kimi K2 Thinking Turbo better than o1?

Kimi K2 Thinking Turbo is the stronger model overall, scoring 45.8 to 40.9 on the Noometry Index.

Is Kimi K2 Thinking Turbo or o1 better for coding?

o1 scores higher on coding benchmarks: 46.1 versus 38.4 in the Noometry coding category.

How many benchmarks do Kimi K2 Thinking Turbo and o1 share?

20 benchmarks have published results for both models. Kimi K2 Thinking Turbo has 21 scored results on Noometry and o1 has 52.

Related comparisons

Go deeper