Model comparison

Claude 3.5 Haiku vs Kimi K2.5 Instant

Kimi K2.5 Instant is the stronger model overall, scoring 43.6 to 29.2 on the Noometry Index.

Last verified . 17 shared benchmarks.

Claude 3.5 Haiku Anthropic

29.2

Rank #315 Confirmed

Kimi K2.5 Instant Moonshot AI

43.6

Rank #89 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Claude 3.5 Haiku scores higher in 0 categories and Kimi K2.5 Instant in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Kimi K2.5 Instant leads 39.4 to 14.7.
  • Kimi K2.5 Instant has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Haiku and Kimi K2.5 Instant specifications
Claude 3.5 HaikuKimi K2.5 Instant
ProviderAnthropicMoonshot AI
Noometry Index29.243.6
Released2024-10-22—
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4918

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K2.5 Instant leads

Claude 3.5 Haiku: 32.9 (#265), Kimi K2.5 Instant: 42.6 (#97)

Coding benchmarks
BenchmarkClaude 3.5 HaikuKimi K2.5 Instant
LMArena Coding12861484
Aider Polyglot28%—
LMArena WebDev—1404
SciCode27.4%—
WeirdML30.7%—
BigCodeBench Instruct46.1%—
LiveBench Coding51.4%—
BigCodeBench Complete59%—
CadEval32%—

Agentic & Tool Use Not comparable

Claude 3.5 Haiku: 28.0 (#95), Kimi K2.5 Instant: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 HaikuKimi K2.5 Instant
BALROG19.3%—

Reasoning Kimi K2.5 Instant leads

Claude 3.5 Haiku: 17.7 (#290), Kimi K2.5 Instant: 29.7 (#90)

Reasoning benchmarks
BenchmarkClaude 3.5 HaikuKimi K2.5 Instant
LMArena Hard Prompts12511443
CritPt0%—
LiveBench Reasoning28.1%—
DTBench56.7%—
LiveBench Data Analysis48.5%—
Epoch Capabilities Index127.15—
LiveBench43.5%—

Math Kimi K2.5 Instant leads

Claude 3.5 Haiku: 14.7 (#300), Kimi K2.5 Instant: 39.4 (#105)

Math benchmarks
BenchmarkClaude 3.5 HaikuKimi K2.5 Instant
LMArena Math12441442
OTIS Mock AIME 2024-20254.3%—
Omni-MATH22.4%—
LiveBench Math35.5%—
MATH Level 546.4%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Kimi K2.5 Instant leads

Claude 3.5 Haiku: 18.7 (#281), Kimi K2.5 Instant: 40.2 (#123)

Knowledge benchmarks
BenchmarkClaude 3.5 HaikuKimi K2.5 Instant
LMArena Expert12081440
GPQA Diamond38.1%—
MMLU-Pro60.5%—
Confabulations36.7%—
GPQA (HELM)36.3%—
MMLU74.3%—

Multimodal Kimi K2.5 Instant leads

Claude 3.5 Haiku: 26.8 (#117), Kimi K2.5 Instant: 40.2 (#50)

Multimodal benchmarks
BenchmarkClaude 3.5 HaikuKimi K2.5 Instant
LMArena Vision10921254
GeoBench34%—

Multilingual Kimi K2.5 Instant leads

Claude 3.5 Haiku: 40.0 (#218), Kimi K2.5 Instant: 52.0 (#94)

Multilingual benchmarks
BenchmarkClaude 3.5 HaikuKimi K2.5 Instant
LMArena Non-English12381406
LMArena Chinese12291449
LMArena French12641403
LMArena German12371413
LMArena Korean11731378
LMArena Russian12531404
LMArena Spanish12611447
LMArena Japanese1175—

Instruction Following Kimi K2.5 Instant leads

Claude 3.5 Haiku: 62.9 (#234), Kimi K2.5 Instant: 75.3 (#65)

Instruction Following benchmarks
BenchmarkClaude 3.5 HaikuKimi K2.5 Instant
LMArena Instruction Following12411430
LiveBench Instruction Following61.9%—
IFEval79.2%—

Long Context Kimi K2.5 Instant leads

Claude 3.5 Haiku: 38.3 (#200), Kimi K2.5 Instant: 43.9 (#83)

Long Context benchmarks
BenchmarkClaude 3.5 HaikuKimi K2.5 Instant
LMArena Longer Query12611435

Writing & Preference Kimi K2.5 Instant leads

Claude 3.5 Haiku: 42.7 (#234), Kimi K2.5 Instant: 60.6 (#95)

Writing & Preference benchmarks
BenchmarkClaude 3.5 HaikuKimi K2.5 Instant
LMArena Text12551420
LMArena Creative Writing12331381
LMArena Multi-Turn12651427
Short-Story Creative Writing73.5%—
EQ-Bench Creative Writing1146—
WildBench76%—
LiveBench Language35.4%—

Frequently asked questions

Is Claude 3.5 Haiku better than Kimi K2.5 Instant?

Kimi K2.5 Instant is the stronger model overall, scoring 43.6 to 29.2 on the Noometry Index.

Is Claude 3.5 Haiku or Kimi K2.5 Instant better for coding?

Kimi K2.5 Instant scores higher on coding benchmarks: 42.6 versus 32.9 in the Noometry coding category.

How many benchmarks do Claude 3.5 Haiku and Kimi K2.5 Instant share?

17 benchmarks have published results for both models. Claude 3.5 Haiku has 49 scored results on Noometry and Kimi K2.5 Instant has 18.

Related comparisons

Go deeper