Model comparison

Kimi K2 (Jul 2025) vs Llama 3.1-405B

Kimi K2 (Jul 2025) is the stronger model overall, scoring 41.2 to 30.7 on the Noometry Index.

Last verified . 29 shared benchmarks.

Kimi K2 (Jul 2025) Moonshot AI

41.2

Rank #140 Confirmed

Llama 3.1-405B Meta

30.7

Rank #288 Confirmed

Summary

  • They share 29 benchmarks with published results for both. Kimi K2 (Jul 2025) scores higher in 9 categories and Llama 3.1-405B in 0 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Kimi K2 (Jul 2025) leads 42.7 to 18.4.
  • The biggest single-benchmark swing is Omni-MATH: 65.4% for Kimi K2 (Jul 2025) and 24.9% for Llama 3.1-405B.

Side by side

Kimi K2 (Jul 2025) and Llama 3.1-405B specifications
Kimi K2 (Jul 2025)Llama 3.1-405B
ProviderMoonshot AIMeta
Noometry Index41.230.7
Released2025-07-122024-07-23
WeightsOpenOpen
Context window262K—
Max output262K—
Input $ / M tokens$0.57—
Output $ / M tokens$2.30—
Results tracked4242

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2 (Jul 2025) leads

Kimi K2 (Jul 2025): 42.4 (#102), Llama 3.1-405B: 33.1 (#262)

Coding benchmarks
BenchmarkKimi K2 (Jul 2025)Llama 3.1-405B
WeirdML42.8%21.4%
LMArena Coding13991291
SWE-bench Verified (bash only)63.4%—
Aider Polyglot59.1%—
GSO4.9%—
ALE-Bench597.5—

Agentic & Tool Use Kimi K2 (Jul 2025) leads

Kimi K2 (Jul 2025): 32.4 (#64), Llama 3.1-405B: 21.0 (#140)

Agentic & Tool Use benchmarks
BenchmarkKimi K2 (Jul 2025)Llama 3.1-405B
Terminal-Bench35.7%—
Berkeley Function Calling Leaderboard59.1%—
TheAgentCompany—7.4%
Cybench—7.5%
METR Time Horizons59.2%—

Reasoning Kimi K2 (Jul 2025) leads

Kimi K2 (Jul 2025): 23.3 (#179), Llama 3.1-405B: 16.8 (#300)

Reasoning benchmarks
BenchmarkKimi K2 (Jul 2025)Llama 3.1-405B
SimpleBench26.3%23%
Kagi LLM Benchmark64.4%45%
LMArena Hard Prompts13841269
Epoch Capabilities Index146.01128.75
ForecastBench60.259.9
DTBench—61.4%
BIG-Bench Hard—82.9%
HellaSwag—89.2%
PIQA—85.9%
WinoGrande—89.2%

Math Kimi K2 (Jul 2025) leads

Kimi K2 (Jul 2025): 42.7 (#83), Llama 3.1-405B: 18.4 (#290)

Math benchmarks
BenchmarkKimi K2 (Jul 2025)Llama 3.1-405B
Omni-MATH65.4%24.9%
LMArena Math13971281
OTIS Mock AIME 2024-2025—9.7%
MATH Level 5—49.8%
FrontierMath (Feb 2025 set)21.4%—
FrontierMath Tier 4 (v1)0%—

Knowledge Kimi K2 (Jul 2025) leads

Kimi K2 (Jul 2025): 37.3 (#157), Llama 3.1-405B: 30.4 (#227)

Knowledge benchmarks
BenchmarkKimi K2 (Jul 2025)Llama 3.1-405B
MMLU-Pro81.9%72.3%
Confabulations20.4%17.6%
GPQA (HELM)65.3%52.2%
LMArena Expert13651243
GPQA Diamond—50.9%
Vectara Hallucination Rate17.9%—
ARC (AI2) Challenge—95.3%
MMLU—84.5%
TriviaQA—82.7%

Multilingual Kimi K2 (Jul 2025) leads

Kimi K2 (Jul 2025): 49.6 (#130), Llama 3.1-405B: 40.7 (#214)

Multilingual benchmarks
BenchmarkKimi K2 (Jul 2025)Llama 3.1-405B
LMArena Non-English13721248
LMArena Chinese14151242
LMArena French13791279
LMArena German13871252
LMArena Japanese13491208
LMArena Korean13251184
LMArena Russian13851265
LMArena Spanish13861260

Instruction Following Kimi K2 (Jul 2025) leads

Kimi K2 (Jul 2025): 71.1 (#156), Llama 3.1-405B: 65.9 (#214)

Instruction Following benchmarks
BenchmarkKimi K2 (Jul 2025)Llama 3.1-405B
IFEval85%81.1%
LMArena Instruction Following13481259

Long Context Kimi K2 (Jul 2025) leads

Kimi K2 (Jul 2025): 41.2 (#145), Llama 3.1-405B: 38.4 (#197)

Long Context benchmarks
BenchmarkKimi K2 (Jul 2025)Llama 3.1-405B
LMArena Longer Query13531266
Fiction.LiveBench66.7%—
CL-bench17.6%—

Writing & Preference Kimi K2 (Jul 2025) leads

Kimi K2 (Jul 2025): 62.3 (#78), Llama 3.1-405B: 38.9 (#251)

Writing & Preference benchmarks
BenchmarkKimi K2 (Jul 2025)Llama 3.1-405B
LMArena Text13801284
LMArena Creative Writing13501262
EQ-Bench Creative Writing1666870
WildBench86.2%78.3%
LMArena Multi-Turn13711297
Short-Story Creative Writing85.6%—

Frequently asked questions

Is Kimi K2 (Jul 2025) better than Llama 3.1-405B?

Kimi K2 (Jul 2025) is the stronger model overall, scoring 41.2 to 30.7 on the Noometry Index.

Is Kimi K2 (Jul 2025) or Llama 3.1-405B better for coding?

Kimi K2 (Jul 2025) scores higher on coding benchmarks: 42.4 versus 33.1 in the Noometry coding category.

How many benchmarks do Kimi K2 (Jul 2025) and Llama 3.1-405B share?

29 benchmarks have published results for both models. Kimi K2 (Jul 2025) has 42 scored results on Noometry and Llama 3.1-405B has 42.

Related comparisons

Go deeper