Model comparison

Hy3 vs Kimi K2 (Jul 2025)

Hy3 is the stronger model overall, scoring 44.2 to 41.2 on the Noometry Index.

Last verified . 17 shared benchmarks.

Hy3 Tencent

44.2

Rank #79 Confirmed

Kimi K2 (Jul 2025) Moonshot AI

41.2

Rank #140 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Hy3 scores higher in 6 categories and Kimi K2 (Jul 2025) in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Hy3 leads 46.8 to 42.4.
  • Hy3 is cheaper at $0.0825 / $0.33 per million input/output tokens, against $0.57 / $2.30 for Kimi K2 (Jul 2025).

Side by side

Hy3 and Kimi K2 (Jul 2025) specifications
Hy3Kimi K2 (Jul 2025)
ProviderTencentMoonshot AI
Noometry Index44.241.2
Released2026-07-062025-07-12
WeightsOpenOpen
Context window262K262K
Max output128K262K
Input $ / M tokens$0.0825$0.57
Output $ / M tokens$0.33$2.30
Results tracked1942

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy3 leads

Hy3: 46.8 (#63), Kimi K2 (Jul 2025): 42.4 (#102)

Coding benchmarks
BenchmarkHy3Kimi K2 (Jul 2025)
LMArena Coding14641399
SWE-bench Verified (bash only)—63.4%
Aider Polyglot—59.1%
LMArena WebDev1508—
GSO—4.9%
WeirdML—42.8%
ALE-Bench—597.5

Agentic & Tool Use Not comparable

Hy3: —, Kimi K2 (Jul 2025): 32.4 (#64)

Agentic & Tool Use benchmarks
BenchmarkHy3Kimi K2 (Jul 2025)
Terminal-Bench—35.7%
Berkeley Function Calling Leaderboard—59.1%
METR Time Horizons—59.2%

Reasoning Hy3 leads

Hy3: 26.1 (#136), Kimi K2 (Jul 2025): 23.3 (#179)

Reasoning benchmarks
BenchmarkHy3Kimi K2 (Jul 2025)
LMArena Hard Prompts14471384
SimpleBench—26.3%
Kagi LLM Benchmark—64.4%
NYT Connections (extended)41.2%—
Epoch Capabilities Index—146.01
ForecastBench—60.2

Math Kimi K2 (Jul 2025) leads

Hy3: 40.1 (#93), Kimi K2 (Jul 2025): 42.7 (#83)

Math benchmarks
BenchmarkHy3Kimi K2 (Jul 2025)
LMArena Math14751397
Omni-MATH—65.4%
FrontierMath (Feb 2025 set)—21.4%
FrontierMath Tier 4 (v1)—0%

Knowledge Hy3 leads

Hy3: 40.8 (#114), Kimi K2 (Jul 2025): 37.3 (#157)

Knowledge benchmarks
BenchmarkHy3Kimi K2 (Jul 2025)
LMArena Expert14601365
MMLU-Pro—81.9%
Confabulations—20.4%
Vectara Hallucination Rate—17.9%
GPQA (HELM)—65.3%

Multilingual Hy3 leads

Hy3: 53.5 (#65), Kimi K2 (Jul 2025): 49.6 (#130)

Multilingual benchmarks
BenchmarkHy3Kimi K2 (Jul 2025)
LMArena Non-English14261372
LMArena Chinese14931415
LMArena French14611379
LMArena German14391387
LMArena Japanese13921349
LMArena Korean13951325
LMArena Russian14321385
LMArena Spanish14561386

Instruction Following Hy3 leads

Hy3: 75.1 (#70), Kimi K2 (Jul 2025): 71.1 (#156)

Instruction Following benchmarks
BenchmarkHy3Kimi K2 (Jul 2025)
LMArena Instruction Following14261348
IFEval—85%

Long Context Hy3 leads

Hy3: 44.1 (#75), Kimi K2 (Jul 2025): 41.2 (#145)

Long Context benchmarks
BenchmarkHy3Kimi K2 (Jul 2025)
LMArena Longer Query14421353
Fiction.LiveBench—66.7%
CL-bench—17.6%

Writing & Preference Too close to call

Hy3: 62.2 (#81), Kimi K2 (Jul 2025): 62.3 (#78)

Writing & Preference benchmarks
BenchmarkHy3Kimi K2 (Jul 2025)
LMArena Text14391380
LMArena Creative Writing14021350
LMArena Multi-Turn14361371
Short-Story Creative Writing—85.6%
EQ-Bench Creative Writing—1666
WildBench—86.2%

Frequently asked questions

Is Hy3 better than Kimi K2 (Jul 2025)?

Hy3 is the stronger model overall, scoring 44.2 to 41.2 on the Noometry Index.

Which is cheaper, Hy3 or Kimi K2 (Jul 2025)?

Hy3 is cheaper. It lists at $0.0825 per million input tokens and $0.33 per million output tokens; Kimi K2 (Jul 2025) lists at $0.57 and $2.30.

Is Hy3 or Kimi K2 (Jul 2025) better for coding?

Hy3 scores higher on coding benchmarks: 46.8 versus 42.4 in the Noometry coding category.

Which has the bigger context window?

Both accept 262K tokens.

How many benchmarks do Hy3 and Kimi K2 (Jul 2025) share?

17 benchmarks have published results for both models. Hy3 has 19 scored results on Noometry and Kimi K2 (Jul 2025) has 42.

Related comparisons

Go deeper