Model comparison

Kimi K2 Thinking Turbo vs o3

o3 is the stronger model overall, scoring 47.5 to 45.8 on the Noometry Index.

Last verified . 20 shared benchmarks.

Kimi K2 Thinking Turbo Moonshot AI

45.8

Rank #70 Confirmed

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 20 benchmarks with published results for both. Kimi K2 Thinking Turbo scores higher in 2 categories and o3 in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in long context, where o3 leads 53.3 to 43.2.
  • The biggest single-benchmark swing is Chess Puzzles: 20% for Kimi K2 Thinking Turbo and 38% for o3.
  • Kimi K2 Thinking Turbo has downloadable open weights; the other is API-only.

Side by side

Kimi K2 Thinking Turbo and o3 specifications
Kimi K2 Thinking Turboo3
ProviderMoonshot AIOpenAI
Noometry Index45.847.5
Released2025-11-062025-04-16
WeightsOpenProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$2
Output $ / M tokens—$8
Results tracked2163

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3 leads

Kimi K2 Thinking Turbo: 38.4 (#178), o3: 46.8 (#64)

Coding benchmarks
BenchmarkKimi K2 Thinking Turboo3
LMArena Coding14541408
SWE-bench Verified—62.3%
SWE-bench Verified (bash only)—58.4%
Aider Polyglot—81.3%
LMArena WebDev1322—
GSO—8.8%
WeirdML—52.4%
CadEval—74%
ALE-Bench—933.55

Agentic & Tool Use Not comparable

Kimi K2 Thinking Turbo: —, o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkKimi K2 Thinking Turboo3
Berkeley Function Calling Leaderboard—63%
GDPval—30.8%
DeepResearch Bench—45.2%
OSWorld—23%
LMArena Search—1144
METR Time Horizons—65.4%

Reasoning Too close to call

Kimi K2 Thinking Turbo: 32.5 (#75), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkKimi K2 Thinking Turboo3
Chess Puzzles20%38%
LMArena Hard Prompts14281402
ARC-AGI-2—6.5%
SimpleBench—53.1%
Kagi LLM Benchmark—67.6%
ARC-AGI-1—60.8%
CritPt—1.4%
EnigmaEval—13.1%
Mystery Game Puzzles—29%
DTBench—84.8%
LMCA—39.7%
Epoch Capabilities Index—146.86
ForecastBench—62.5

Math o3 leads

Kimi K2 Thinking Turbo: 47.4 (#68), o3: 50.2 (#58)

Math benchmarks
BenchmarkKimi K2 Thinking Turboo3
OTIS Mock AIME 2024-202583.1%84.4%
LMArena Math14291426
FrontierMath (Tiers 1-3)—33.3%
Omni-MATH—71.4%
MATH Level 5—97.8%
FrontierMath (Feb 2025 set)—18.7%
FrontierMath Tier 4 (v1)—2.1%

Knowledge o3 leads

Kimi K2 Thinking Turbo: 50.9 (#69), o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkKimi K2 Thinking Turboo3
GPQA Diamond84.2%81.8%
LMArena Expert14391402
Humanity's Last Exam—20.3%
SimpleQA Verified—49.4%
MMLU-Pro—85.9%
Confabulations—14.4%
GPQA (HELM)—75.3%

Multimodal Not comparable

Kimi K2 Thinking Turbo: —, o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkKimi K2 Thinking Turboo3
LMArena Vision—1214
GeoBench—74%
VPCT—52%

Multilingual Too close to call

Kimi K2 Thinking Turbo: 51.4 (#109), o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkKimi K2 Thinking Turboo3
LMArena Non-English13981401
LMArena Chinese14561437
LMArena French14251430
LMArena German13901420
LMArena Japanese13571403
LMArena Korean13301370
LMArena Russian13911406
LMArena Spanish14061395

Instruction Following Kimi K2 Thinking Turbo leads

Kimi K2 Thinking Turbo: 74.0 (#109), o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkKimi K2 Thinking Turboo3
LMArena Instruction Following14031368
IFEval—86.9%

Long Context o3 leads

Kimi K2 Thinking Turbo: 43.2 (#102), o3: 53.3 (#6)

Long Context benchmarks
BenchmarkKimi K2 Thinking Turboo3
LMArena Longer Query14151372
Fiction.LiveBench—88.9%
CL-bench—17.8%

Writing & Preference o3 leads

Kimi K2 Thinking Turbo: 60.0 (#104), o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkKimi K2 Thinking Turboo3
LMArena Text14151410
LMArena Creative Writing13741359
LMArena Multi-Turn14141405
Short-Story Creative Writing—83.9%
EQ-Bench Creative Writing—1676
WildBench—86.1%

Frequently asked questions

Is Kimi K2 Thinking Turbo better than o3?

o3 is the stronger model overall, scoring 47.5 to 45.8 on the Noometry Index.

Is Kimi K2 Thinking Turbo or o3 better for coding?

o3 scores higher on coding benchmarks: 46.8 versus 38.4 in the Noometry coding category.

How many benchmarks do Kimi K2 Thinking Turbo and o3 share?

20 benchmarks have published results for both models. Kimi K2 Thinking Turbo has 21 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper