Model comparison

Kimi K2.6 vs o3-pro

Kimi K2.6 is the stronger model overall, scoring 47.7 to 42.9 on the Noometry Index.

Last verified . 5 shared benchmarks.

Kimi K2.6 Moonshot AI

47.7

Rank #60 Confirmed

o3-pro OpenAI

42.9

Rank #105 Confirmed

Summary

  • They share 5 benchmarks with published results for both. Kimi K2.6 scores higher in 3 categories and o3-pro in 2 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in long context, where o3-pro leads 72.2 to 44.9.
  • The biggest single-benchmark swing is Vectara Hallucination Rate: 10.8% for Kimi K2.6 and 23.3% for o3-pro.
  • Kimi K2.6 is cheaper at $0.95 / $4 per million input/output tokens, against $20 / $80 for o3-pro.
  • Kimi K2.6 accepts more context: 262K tokens versus 200K.
  • Kimi K2.6 has downloadable open weights; the other is API-only.

Side by side

Kimi K2.6 and o3-pro specifications
Kimi K2.6o3-pro
ProviderMoonshot AIOpenAI
Noometry Index47.742.9
Released2026-04-202025-06-10
WeightsOpenProprietary
Context window262K200K
Max output262K100K
Input $ / M tokens$0.95$20
Output $ / M tokens$4$80
Results tracked5112

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3-pro leads

Kimi K2.6: 50.7 (#43), o3-pro: 55.5 (#24)

Coding benchmarks
BenchmarkKimi K2.6o3-pro
WeirdML55.9%58.2%
SWE-bench Verified76.7%—
Aider Polyglot—84.9%
LMArena WebDev1509—
SciCode53.5%—
LMArena Coding1488—
ALE-Bench1,093—

Agentic & Tool Use Not comparable

Kimi K2.6: 21.9 (#137), o3-pro: —

Agentic & Tool Use benchmarks
BenchmarkKimi K2.6o3-pro
OSWorld 2.04.6%—
ExploitBench18.4%—
GBAEval0.9%—
GDP.pdf12%—
Vending-Bench 26,205—

Reasoning Kimi K2.6 leads

Kimi K2.6: 40.5 (#55), o3-pro: 23.8 (#171)

Reasoning benchmarks
BenchmarkKimi K2.6o3-pro
DTBench90.9%86.9%
LMCA37.3%38.5%
Epoch Capabilities Index151.05147.42
ARC-AGI-2—4.9%
Kagi LLM Benchmark—72.1%
NYT Connections (extended)87.2%—
ARC-AGI-1—59.3%
CritPt8%—
Chess Puzzles26%—
EBR-Bench2.4%—
LMArena Hard Prompts1470—
Mystery Game Puzzles18%—

Math Not comparable

Kimi K2.6: 57.0 (#41), o3-pro: —

Knowledge Kimi K2.6 leads

Kimi K2.6: 54.0 (#54), o3-pro: 29.5 (#238)

Knowledge benchmarks
BenchmarkKimi K2.6o3-pro
Vectara Hallucination Rate10.8%23.3%
GPQA Diamond90.8%—
SimpleQA Verified34.9%—
Confabulations—14.2%
LMArena Expert1491—

Multimodal Not comparable

Kimi K2.6: 31.6 (#103), o3-pro: —

Multimodal benchmarks
BenchmarkKimi K2.6o3-pro
LMArena Vision1283—
Blueprint-Bench 23.9%—
Furniture Assembly21.7%—
LMArena Document1451—

Multilingual Not comparable

Kimi K2.6: 54.9 (#37), o3-pro: —

Multilingual benchmarks
BenchmarkKimi K2.6o3-pro
LMArena Non-English1446—
LMArena Chinese1521—
LMArena French1471—
LMArena German1450—
LMArena Japanese1443—
LMArena Korean1427—
LMArena Russian1446—
LMArena Spanish1464—

Instruction Following Not comparable

Kimi K2.6: 76.3 (#43), o3-pro: —

Instruction Following benchmarks
BenchmarkKimi K2.6o3-pro
LMArena Instruction Following1451—

Long Context o3-pro leads

Kimi K2.6: 44.9 (#52), o3-pro: 72.2 (#1)

Long Context benchmarks
BenchmarkKimi K2.6o3-pro
Fiction.LiveBench—97.2%
LMArena Longer Query1468—

Writing & Preference Kimi K2.6 leads

Kimi K2.6: 68.5 (#26), o3-pro: 57.1 (#133)

Writing & Preference benchmarks
BenchmarkKimi K2.6o3-pro
LMArena Text1455—
LMArena Creative Writing1434—
Short-Story Creative Writing—84.4%
EQ-Bench Creative Writing1725—
EQ-Bench 41202—
LMArena Multi-Turn1453—

Frequently asked questions

Is Kimi K2.6 better than o3-pro?

Kimi K2.6 is the stronger model overall, scoring 47.7 to 42.9 on the Noometry Index.

Which is cheaper, Kimi K2.6 or o3-pro?

Kimi K2.6 is cheaper. It lists at $0.95 per million input tokens and $4 per million output tokens; o3-pro lists at $20 and $80.

Is Kimi K2.6 or o3-pro better for coding?

o3-pro scores higher on coding benchmarks: 55.5 versus 50.7 in the Noometry coding category.

Which has the bigger context window?

Kimi K2.6 does, with 262K tokens against 200K.

How many benchmarks do Kimi K2.6 and o3-pro share?

5 benchmarks have published results for both models. Kimi K2.6 has 51 scored results on Noometry and o3-pro has 12.

Related comparisons

Go deeper