Model comparison

Hy4 preview vs Kimi K2.5

Kimi K2.5 is the stronger model overall, scoring 48.1 to 45.3 on the Noometry Index.

Last verified . 2 shared benchmarks.

Hy4 preview Tencent

45.3

Rank #73 Reported

Kimi K2.5 Moonshot AI

48.1

Rank #57 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Hy4 preview scores higher in 3 categories and Kimi K2.5 in 0 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in math, where Hy4 preview leads 55.7 to 51.8.
  • Kimi K2.5 is cheaper at $0.45 / $2.25 per million input/output tokens, against $0.75 / $2.25 for Hy4 preview.
  • Hy4 preview accepts more context: 1.05M tokens versus 262K.

Side by side

Hy4 preview and Kimi K2.5 specifications
Hy4 previewKimi K2.5
ProviderTencentMoonshot AI
Noometry Index45.348.1
Released2026-08-282026-01-27
WeightsOpenOpen
Context window1.05M262K
Max output64K262K
Input $ / M tokens$0.75$0.45
Output $ / M tokens$2.25$2.25
Results tracked351

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy4 preview leads

Hy4 preview: 51.6 (#38), Kimi K2.5: 48.8 (#53)

Coding benchmarks
BenchmarkHy4 previewKimi K2.5
LMArena WebDev16321437
SWE-bench Verified—73.8%
SWE-bench Verified (bash only)—70.8%
SWE-bench Multilingual—67.3%
SciCode—49%
WeirdML—45.6%
LMArena Coding—1474
ALE-Bench—821.65

Agentic & Tool Use Not comparable

Hy4 preview: —, Kimi K2.5: 34.2 (#48)

Agentic & Tool Use benchmarks
BenchmarkHy4 previewKimi K2.5
Terminal-Bench—43.2%
OSWorld—63.3%
Vending-Bench 2—1,198

Reasoning Too close to call

Hy4 preview: 31.9 (#79), Kimi K2.5: 31.2 (#80)

Reasoning benchmarks
BenchmarkHy4 previewKimi K2.5
NYT Connections (extended)68.2%69.9%
ARC-AGI-2—11.8%
SimpleBench—46.8%
Kagi LLM Benchmark—78.5%
ARC-AGI-1—65.3%
CritPt—3.1%
Chess Puzzles—12%
EnigmaEval—3.4%
Thematic Generalization—69.4%
LMArena Hard Prompts—1453
Epoch Capabilities Index—148.03

Math Hy4 preview leads

Hy4 preview: 55.7 (#42), Kimi K2.5: 51.8 (#53)

Math benchmarks
BenchmarkHy4 previewKimi K2.5
MathArena Final-Answer Competitions—62.3%
OTIS Mock AIME 2024-2025—92.2%
ProofBench75%—
LMArena Math—1470
FrontierMath (Feb 2025 set)—27.9%
FrontierMath Tier 4 (v1)—4.2%

Knowledge Not comparable

Hy4 preview: —, Kimi K2.5: 53.6 (#56)

Knowledge benchmarks
BenchmarkHy4 previewKimi K2.5
GPQA Diamond—87.6%
Humanity's Last Exam—24.4%
SimpleQA Verified—34.3%
Vectara Hallucination Rate—14.2%
LMArena Expert—1466

Multimodal Not comparable

Hy4 preview: —, Kimi K2.5: 41.1 (#39)

Multimodal benchmarks
BenchmarkHy4 previewKimi K2.5
LMArena Vision—1269
LMArena Document—1430

Multilingual Not comparable

Hy4 preview: —, Kimi K2.5: 53.9 (#53)

Multilingual benchmarks
BenchmarkHy4 previewKimi K2.5
LMArena Non-English—1433
LMArena Chinese—1495
LMArena French—1454
LMArena German—1441
LMArena Japanese—1421
LMArena Korean—1410
LMArena Russian—1435
LMArena Spanish—1450

Instruction Following Not comparable

Hy4 preview: —, Kimi K2.5: 75.3 (#64)

Instruction Following benchmarks
BenchmarkHy4 previewKimi K2.5
LMArena Instruction Following—1431

Long Context Not comparable

Hy4 preview: —, Kimi K2.5: 52.1 (#7)

Long Context benchmarks
BenchmarkHy4 previewKimi K2.5
Fiction.LiveBench—86.1%
CL-bench—19.3%
CL-bench Life—13.2%
LMArena Longer Query—1445

Writing & Preference Not comparable

Hy4 preview: —, Kimi K2.5: 65.1 (#53)

Writing & Preference benchmarks
BenchmarkHy4 previewKimi K2.5
LMArena Text—1445
LMArena Creative Writing—1423
EQ-Bench Creative Writing—1579
LMArena Multi-Turn—1444

Frequently asked questions

Is Hy4 preview better than Kimi K2.5?

Kimi K2.5 is the stronger model overall, scoring 48.1 to 45.3 on the Noometry Index.

Which is cheaper, Hy4 preview or Kimi K2.5?

Kimi K2.5 is cheaper. It lists at $0.45 per million input tokens and $2.25 per million output tokens; Hy4 preview lists at $0.75 and $2.25.

Is Hy4 preview or Kimi K2.5 better for coding?

Hy4 preview scores higher on coding benchmarks: 51.6 versus 48.8 in the Noometry coding category.

Which has the bigger context window?

Hy4 preview does, with 1.05M tokens against 262K.

How many benchmarks do Hy4 preview and Kimi K2.5 share?

2 benchmarks have published results for both models. Hy4 preview has 3 scored results on Noometry and Kimi K2.5 has 51.

Related comparisons

Go deeper