Model comparison

Hy4 preview vs Kimi K2 (Jul 2025)

Hy4 preview is the stronger model overall, scoring 45.3 to 41.2 on the Noometry Index.

Last verified . 0 shared benchmarks.

Hy4 preview Tencent

45.3

Rank #73 Reported

Kimi K2 (Jul 2025) Moonshot AI

41.2

Rank #140 Confirmed

Summary

  • The widest gap is in math, where Hy4 preview leads 55.7 to 42.7.
  • Kimi K2 (Jul 2025) is cheaper at $0.57 / $2.30 per million input/output tokens, against $0.75 / $2.25 for Hy4 preview.
  • Hy4 preview accepts more context: 1.05M tokens versus 262K.

Side by side

Hy4 preview and Kimi K2 (Jul 2025) specifications
Hy4 previewKimi K2 (Jul 2025)
ProviderTencentMoonshot AI
Noometry Index45.341.2
Released2026-08-282025-07-12
WeightsOpenOpen
Context window1.05M262K
Max output64K262K
Input $ / M tokens$0.75$0.57
Output $ / M tokens$2.25$2.30
Results tracked342

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy4 preview leads

Hy4 preview: 51.6 (#38), Kimi K2 (Jul 2025): 42.4 (#102)

Coding benchmarks
BenchmarkHy4 previewKimi K2 (Jul 2025)
SWE-bench Verified (bash only)—63.4%
Aider Polyglot—59.1%
LMArena WebDev1632—
GSO—4.9%
WeirdML—42.8%
LMArena Coding—1399
ALE-Bench—597.5

Agentic & Tool Use Not comparable

Hy4 preview: —, Kimi K2 (Jul 2025): 32.4 (#64)

Agentic & Tool Use benchmarks
BenchmarkHy4 previewKimi K2 (Jul 2025)
Terminal-Bench—35.7%
Berkeley Function Calling Leaderboard—59.1%
METR Time Horizons—59.2%

Reasoning Hy4 preview leads

Hy4 preview: 31.9 (#79), Kimi K2 (Jul 2025): 23.3 (#179)

Reasoning benchmarks
BenchmarkHy4 previewKimi K2 (Jul 2025)
SimpleBench—26.3%
Kagi LLM Benchmark—64.4%
NYT Connections (extended)68.2%—
LMArena Hard Prompts—1384
Epoch Capabilities Index—146.01
ForecastBench—60.2

Math Hy4 preview leads

Hy4 preview: 55.7 (#42), Kimi K2 (Jul 2025): 42.7 (#83)

Math benchmarks
BenchmarkHy4 previewKimi K2 (Jul 2025)
ProofBench75%—
Omni-MATH—65.4%
LMArena Math—1397
FrontierMath (Feb 2025 set)—21.4%
FrontierMath Tier 4 (v1)—0%

Knowledge Not comparable

Hy4 preview: —, Kimi K2 (Jul 2025): 37.3 (#157)

Knowledge benchmarks
BenchmarkHy4 previewKimi K2 (Jul 2025)
MMLU-Pro—81.9%
Confabulations—20.4%
Vectara Hallucination Rate—17.9%
GPQA (HELM)—65.3%
LMArena Expert—1365

Multilingual Not comparable

Hy4 preview: —, Kimi K2 (Jul 2025): 49.6 (#130)

Multilingual benchmarks
BenchmarkHy4 previewKimi K2 (Jul 2025)
LMArena Non-English—1372
LMArena Chinese—1415
LMArena French—1379
LMArena German—1387
LMArena Japanese—1349
LMArena Korean—1325
LMArena Russian—1385
LMArena Spanish—1386

Instruction Following Not comparable

Hy4 preview: —, Kimi K2 (Jul 2025): 71.1 (#156)

Instruction Following benchmarks
BenchmarkHy4 previewKimi K2 (Jul 2025)
IFEval—85%
LMArena Instruction Following—1348

Long Context Not comparable

Hy4 preview: —, Kimi K2 (Jul 2025): 41.2 (#145)

Long Context benchmarks
BenchmarkHy4 previewKimi K2 (Jul 2025)
Fiction.LiveBench—66.7%
CL-bench—17.6%
LMArena Longer Query—1353

Writing & Preference Not comparable

Hy4 preview: —, Kimi K2 (Jul 2025): 62.3 (#78)

Writing & Preference benchmarks
BenchmarkHy4 previewKimi K2 (Jul 2025)
LMArena Text—1380
LMArena Creative Writing—1350
Short-Story Creative Writing—85.6%
EQ-Bench Creative Writing—1666
WildBench—86.2%
LMArena Multi-Turn—1371

Frequently asked questions

Is Hy4 preview better than Kimi K2 (Jul 2025)?

Hy4 preview is the stronger model overall, scoring 45.3 to 41.2 on the Noometry Index.

Which is cheaper, Hy4 preview or Kimi K2 (Jul 2025)?

Kimi K2 (Jul 2025) is cheaper. It lists at $0.57 per million input tokens and $2.30 per million output tokens; Hy4 preview lists at $0.75 and $2.25.

Is Hy4 preview or Kimi K2 (Jul 2025) better for coding?

Hy4 preview scores higher on coding benchmarks: 51.6 versus 42.4 in the Noometry coding category.

Which has the bigger context window?

Hy4 preview does, with 1.05M tokens against 262K.

How many benchmarks do Hy4 preview and Kimi K2 (Jul 2025) share?

0 benchmarks have published results for both models. Hy4 preview has 3 scored results on Noometry and Kimi K2 (Jul 2025) has 42.

Related comparisons

Go deeper