Moonshot AI, open weights

Kimi K2 (Jul 2025)

Kimi K2 (Jul 2025) by Moonshot AI ranks 140th of 354 ranked models on the Noometry Index as of October 2026, with a score of 41.2. Its strongest category is agentic & tool use, where it ranks 64th. API pricing starts at $0.57 per million input tokens and $2.30 per million output tokens, with a 262K-token context window.

Last verified

Specifications

Noometry rank
#140 of 354
Index score
41.2
Evidence
Confirmed 42 results
Provider
Moonshot AI
Released
July 12, 2025
Weights
Open weights
Reasoning
Yes
Context window
262K
Max output
262K
Input price
$0.57 / M
Output price
$2.30 / M
Blended price
$1 / M
Output speed
201 tokens/s Kagi
Value
#116 of 219
Knowledge cutoff
August 2024
Input
text

Category scores

Each category score combines every public result we have in that category.

Kimi K2 (Jul 2025) category scores
  1. Coding 42.4
  2. Agentic & Tool Use 32.4
  3. Reasoning 23.3
  4. Math 42.7
  5. Knowledge 37.3
  6. Multilingual 49.6
  7. Instruction Following 71.1
  8. Long Context 41.2
  9. Writing & Preference 62.3
Kimi K2 (Jul 2025) category ranks
CategoryScoreRankResults
Coding42.4#1025
Agentic & Tool Use32.4#642
Reasoning23.3#1793
Math42.7#832
Knowledge37.3#1575
Multilingual49.6#1301
Instruction Following71.1#1562
Long Context41.2#1453
Writing & Preference62.3#786

Strengths and weaknesses

Categories where Kimi K2 (Jul 2025) places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Kimi K2 (Jul 2025): strongest categories
CategoryScorevs medianRank
Writing & Preference62.3+8.5#78 of 312, top 25%
Math42.7+6.1#83 of 327, top 26%
Coding42.4+3.7#102 of 340, top 30%

Weakest categories

Kimi K2 (Jul 2025): weakest categories
CategoryScorevs medianRank
Instruction Following71.1−0.1#156 of 305, top 52%
Reasoning23.3−0.3#179 of 350, top 52%
Knowledge37.3−0.0#157 of 314, top 50%

Closest competitors

The models ranked just above and below Kimi K2 (Jul 2025). When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Kimi K2 (Jul 2025)
ModelRankScoreBlended $/MSpeed
Grok 4.1 Fast#13641.4$0.28—Compare
GLM-4.6V#13741.3$0.45—Compare
MiMo-V2-Flash#13841.3$0.18—Compare
Hunyuan Turbos 20250226#13941.3——Compare
Grok-3 mini#14141.2—10Compare
Claude Opus 4.1#14241.0$30—Compare
o1#14340.9$26.25—Compare
Gemini 3.1 Flash Lite#14440.8$0.5610Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Kimi K2 (Jul 2025) Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified (bash only)43.8%SWE-bench2025-08-07
SWE-bench Verified (bash only)63.4%#18 of 39, top 47%SWE-bench2025-12-10
Aider Polyglot59.1%#16 of 44, top 37%Epoch AI
Aider Polyglot59.1%#16 of 44, top 37%Epoch AI
GSO4.9%#22 of 31, top 71%Epoch AI
WeirdML39.4%Epoch AI
WeirdML36.7%Epoch AI
WeirdML42.8%#68 of 119, top 58%Epoch AI
WeirdML39.4%Epoch AI
LMArena Coding1399#137 of 294, top 47%LMArena2026-10-08
LMArena Coding1377LMArena2026-10-08
ALE-Bench597.5#81 of 105, top 78%Epoch AI
ALE-Bench267.12Epoch AI

Agentic & Tool Use

Kimi K2 (Jul 2025) Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Terminal-Bench27.8%Epoch AI
Terminal-Bench35.7%#27 of 41, top 66%Epoch AI
Berkeley Function Calling Leaderboard59.1%#9 of 49, top 19%fcBerkeley Function Calling Leaderboard
METR Time Horizons59.2%#19 of 32, top 60%Epoch AI

Reasoning

Kimi K2 (Jul 2025) Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
SimpleBench26.3%#65 of 77, top 85%Epoch AI
Kagi LLM Benchmark64.4%#33 of 99, top 34%Kagi LLM Benchmark
Kagi LLM Benchmark52.2%Kagi LLM Benchmark
Kagi LLM Benchmark45%Kagi LLM Benchmark
LMArena Hard Prompts1366LMArena2026-10-08
LMArena Hard Prompts1384#136 of 297, top 46%LMArena2026-10-08
Epoch Capabilities Index146.01#76 of 213, top 36%Epoch AI2025-11-06
Epoch Capabilities Index140.13Epoch AI2025-07-12
ForecastBench60.2#30 of 72, top 42%Epoch AI
ForecastBench59.8Epoch AI

Math

Kimi K2 (Jul 2025) Math benchmark results
BenchmarkScorePositionSettingSourceDate
Omni-MATH65.4%#6 of 57, top 11%HELM Capabilities
LMArena Math1397#127 of 285, top 45%LMArena2026-10-08
LMArena Math1367LMArena2026-10-08
FrontierMath (Feb 2025 set)21.4%#28 of 68, top 42%Epoch AI2025-12-05
FrontierMath Tier 4 (v1)0%#52 of 55, top 95%Epoch AI2025-12-04

Knowledge

Kimi K2 (Jul 2025) Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
MMLU-Pro81.9%#13 of 58, top 23%HELM Capabilities
Confabulations (lower is better)20.4%#34 of 51, top 67%Lech Mazur benchmarks
Vectara Hallucination Rate (lower is better)17.9%#88 of 96, top 92%Vectara Hallucination Leaderboard
GPQA (HELM)65.3%#17 of 57, top 30%HELM Capabilities
LMArena Expert1365#143 of 273, top 53%LMArena2026-10-08
LMArena Expert1346LMArena2026-10-08

Multilingual

Kimi K2 (Jul 2025) Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1372#130 of 297, top 44%LMArena2026-10-08
LMArena Non-English1358LMArena2026-10-08
LMArena Chinese1415#131 of 285, top 46%LMArena2026-10-08
LMArena Chinese1401LMArena2026-10-08
LMArena French1373LMArena2026-10-08
LMArena French1379#132 of 223, top 60%LMArena2026-10-08
LMArena German1387#106 of 231, top 46%LMArena2026-10-08
LMArena German1374LMArena2026-10-08
LMArena Japanese1336LMArena2026-10-08
LMArena Japanese1349#98 of 211, top 47%LMArena2026-10-08
LMArena Korean1325#116 of 213, top 55%LMArena2026-10-08
LMArena Korean1287LMArena2026-10-08
LMArena Russian1361LMArena2026-10-08
LMArena Russian1385#126 of 283, top 45%LMArena2026-10-08
LMArena Spanish1386#122 of 226, top 54%LMArena2026-10-08
LMArena Spanish1330LMArena2026-10-08

Instruction Following

Kimi K2 (Jul 2025) Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
IFEval85%#18 of 57, top 32%HELM Capabilities
LMArena Instruction Following1324LMArena2026-10-08
LMArena Instruction Following1348#146 of 298, top 49%LMArena2026-10-08

Long Context

Kimi K2 (Jul 2025) Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
Fiction.LiveBench61.1%Epoch AI
Fiction.LiveBench66.7%#21 of 47, top 45%Epoch AI
CL-bench11.9%Epoch AI
CL-bench17.6%#13 of 19, top 69%Epoch AI
LMArena Longer Query1326LMArena2026-10-08
LMArena Longer Query1353#151 of 291, top 52%LMArena2026-10-08

Writing & Preference

Kimi K2 (Jul 2025) Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1380#135 of 297, top 46%LMArena2026-10-08
LMArena Text1371LMArena2026-10-08
LMArena Creative Writing1324LMArena2026-10-08
LMArena Creative Writing1350#127 of 295, top 44%LMArena2026-10-08
Short-Story Creative Writing85.6%#2 of 39, top 6%Epoch AI
EQ-Bench Creative Writing1631EQ-Bench
EQ-Bench Creative Writing1666#37 of 115, top 33%EQ-Bench
WildBench86.2%#3 of 57, top 6%HELM Capabilities
LMArena Multi-Turn1365LMArena2026-10-08
LMArena Multi-Turn1371#139 of 295, top 48%LMArena2026-10-08

API pricing by provider

Kimi K2 (Jul 2025) API prices
RouteInput $/MOutput $/MCached input $/MChecked
bedrock$0.60$2.50—2026-10-10
openrouter$0.57$2.30—2026-10-10
vertex$0.60$2.50$0.062026-10-10

Compare Kimi K2 (Jul 2025)

Other Moonshot AI models

Frequently asked questions

How good is Kimi K2 (Jul 2025)?

Kimi K2 (Jul 2025) by Moonshot AI ranks 140th of 354 ranked models on the Noometry Index as of October 2026, with a score of 41.2. Its strongest category is agentic & tool use, where it ranks 64th. API pricing starts at $0.57 per million input tokens and $2.30 per million output tokens, with a 262K-token context window.

How much does Kimi K2 (Jul 2025) cost?

Kimi K2 (Jul 2025) costs $0.57 per million input tokens and $2.30 per million output tokens on openrouter.

What is Kimi K2 (Jul 2025)'s context window?

Kimi K2 (Jul 2025) accepts up to 262K tokens of input and can write up to 262K tokens in one response.

Is Kimi K2 (Jul 2025) open source?

Yes. Kimi K2 (Jul 2025)'s weights are downloadable from Hugging Face (moonshotai/Kimi-K2-Instruct); check the license for commercial terms.

How fast is Kimi K2 (Jul 2025)?

Kimi K2 (Jul 2025) generated about 201 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are Kimi K2 (Jul 2025)'s strengths and weaknesses?

Relative to other ranked models, Kimi K2 (Jul 2025) places best in writing & preference, math, coding and lowest in instruction following, reasoning, knowledge.

What is Kimi K2 (Jul 2025) best at?

Its best category is agentic & tool use, where it ranks 64th on Noometry.