Model comparison

Claude 3 Opus vs Kimi K2.5

Kimi K2.5 is the stronger model overall, scoring 48.1 to 29.5 on the Noometry Index.

Last verified . 26 shared benchmarks.

Claude 3 Opus Anthropic

29.5

Rank #310 Confirmed

Kimi K2.5 Moonshot AI

48.1

Rank #57 Confirmed

Summary

  • They share 26 benchmarks with published results for both. Claude 3 Opus scores higher in 0 categories and Kimi K2.5 in 10 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where Kimi K2.5 leads 51.8 to 14.8.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 4.7% for Claude 3 Opus and 92.2% for Kimi K2.5.
  • Kimi K2.5 has downloadable open weights; the other is API-only.

Side by side

Claude 3 Opus and Kimi K2.5 specifications
Claude 3 OpusKimi K2.5
ProviderAnthropicMoonshot AI
Noometry Index29.548.1
Released2024-02-292026-01-27
WeightsProprietaryOpen
Context window—262K
Max output—262K
Input $ / M tokens—$0.45
Output $ / M tokens—$2.25
Results tracked4651

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2.5 leads

Claude 3 Opus: 32.9 (#267), Kimi K2.5: 48.8 (#53)

Coding benchmarks
BenchmarkClaude 3 OpusKimi K2.5
WeirdML19.2%45.6%
LMArena Coding12641474
SWE-bench Verified—73.8%
SWE-bench Verified (bash only)—70.8%
LMArena WebDev—1437
SWE-bench Multilingual—67.3%
SciCode—49%
BigCodeBench Instruct45.5%—
LiveBench Coding38.6%—
BigCodeBench Complete57.4%—
ALE-Bench—821.65
HumanEval+77.4%—
MBPP+73.3%—

Agentic & Tool Use Kimi K2.5 leads

Claude 3 Opus: 24.6 (#116), Kimi K2.5: 34.2 (#48)

Agentic & Tool Use benchmarks
BenchmarkClaude 3 OpusKimi K2.5
Terminal-Bench—43.2%
Cybench10%—
OSWorld—63.3%
METR Time Horizons29.5%—
Vending-Bench 2—1,198

Reasoning Kimi K2.5 leads

Claude 3 Opus: 14.6 (#324), Kimi K2.5: 31.2 (#80)

Reasoning benchmarks
BenchmarkClaude 3 OpusKimi K2.5
SimpleBench23.5%46.8%
Chess Puzzles5%12%
EnigmaEval0.8%3.4%
LMArena Hard Prompts12451453
Epoch Capabilities Index126.91148.03
ARC-AGI-2—11.8%
Kagi LLM Benchmark—78.5%
NYT Connections (extended)—69.9%
ARC-AGI-1—65.3%
CritPt—3.1%
Thematic Generalization—69.4%
LiveBench Reasoning40.6%—
DTBench61.6%—
LiveBench Data Analysis57.9%—
LMCA17%—
ForecastBench58.4—
LiveBench49.2%—
WinoGrande88.5%—

Math Kimi K2.5 leads

Claude 3 Opus: 14.8 (#299), Kimi K2.5: 51.8 (#53)

Math benchmarks
BenchmarkClaude 3 OpusKimi K2.5
OTIS Mock AIME 2024-20254.7%92.2%
LMArena Math12731470
MathArena Final-Answer Competitions—62.3%
LiveBench Math43.6%—
MATH Level 537.5%—
FrontierMath (Feb 2025 set)—27.9%
FrontierMath Tier 4 (v1)—4.2%

Knowledge Kimi K2.5 leads

Claude 3 Opus: 24.5 (#267), Kimi K2.5: 53.6 (#56)

Knowledge benchmarks
BenchmarkClaude 3 OpusKimi K2.5
GPQA Diamond47.2%87.6%
SimpleQA Verified12.6%34.3%
LMArena Expert12231466
Humanity's Last Exam—24.4%
Confabulations22.7%—
Vectara Hallucination Rate—14.2%
MMLU84.6%—

Multimodal Kimi K2.5 leads

Claude 3 Opus: 27.1 (#116), Kimi K2.5: 41.1 (#39)

Multimodal benchmarks
BenchmarkClaude 3 OpusKimi K2.5
LMArena Vision10231269
LMArena Document—1430

Multilingual Kimi K2.5 leads

Claude 3 Opus: 41.4 (#207), Kimi K2.5: 53.9 (#53)

Multilingual benchmarks
BenchmarkClaude 3 OpusKimi K2.5
LMArena Non-English12581433
LMArena Chinese12481495
LMArena French12751454
LMArena German12581441
LMArena Japanese12041421
LMArena Korean11871410
LMArena Russian12801435
LMArena Spanish12461450

Instruction Following Kimi K2.5 leads

Claude 3 Opus: 64.1 (#228), Kimi K2.5: 75.3 (#64)

Instruction Following benchmarks
BenchmarkClaude 3 OpusKimi K2.5
LMArena Instruction Following12481431
LiveBench Instruction Following63.9%—

Long Context Kimi K2.5 leads

Claude 3 Opus: 38.2 (#202), Kimi K2.5: 52.1 (#7)

Long Context benchmarks
BenchmarkClaude 3 OpusKimi K2.5
LMArena Longer Query12591445
Fiction.LiveBench—86.1%
CL-bench—19.3%
CL-bench Life—13.2%

Writing & Preference Kimi K2.5 leads

Claude 3 Opus: 47.2 (#213), Kimi K2.5: 65.1 (#53)

Writing & Preference benchmarks
BenchmarkClaude 3 OpusKimi K2.5
LMArena Text12621445
LMArena Creative Writing12351423
LMArena Multi-Turn12751444
EQ-Bench Creative Writing—1579
LiveBench Language50.4%—

Frequently asked questions

Is Claude 3 Opus better than Kimi K2.5?

Kimi K2.5 is the stronger model overall, scoring 48.1 to 29.5 on the Noometry Index.

Is Claude 3 Opus or Kimi K2.5 better for coding?

Kimi K2.5 scores higher on coding benchmarks: 48.8 versus 32.9 in the Noometry coding category.

How many benchmarks do Claude 3 Opus and Kimi K2.5 share?

26 benchmarks have published results for both models. Claude 3 Opus has 46 scored results on Noometry and Kimi K2.5 has 51.

Related comparisons

Go deeper