Model comparison

Hy3 vs Qwen2.5 Plus 1127

Hy3 is the stronger model overall, scoring 44.2 to 38.8 on the Noometry Index.

Last verified . 14 shared benchmarks.

Hy3 Tencent

44.2

Rank #79 Confirmed

Qwen2.5 Plus 1127 Alibaba (Qwen)

38.8

Rank #181 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Hy3 scores higher in 8 categories and Qwen2.5 Plus 1127 in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Hy3 leads 62.2 to 49.4.
  • Hy3 has downloadable open weights; the other is API-only.

Side by side

Hy3 and Qwen2.5 Plus 1127 specifications
Hy3Qwen2.5 Plus 1127
ProviderTencentAlibaba (Qwen)
Noometry Index44.238.8
Released2026-07-06—
WeightsOpenProprietary
Context window262K—
Max output128K—
Input $ / M tokens$0.13—
Output $ / M tokens$0.53—
Results tracked1914

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy3 leads

Hy3: 46.8 (#63), Qwen2.5 Plus 1127: 38.5 (#175)

Coding benchmarks
BenchmarkHy3Qwen2.5 Plus 1127
LMArena Coding14641314
LMArena WebDev1508—

Reasoning Too close to call

Hy3: 26.1 (#136), Qwen2.5 Plus 1127: 25.9 (#141)

Reasoning benchmarks
BenchmarkHy3Qwen2.5 Plus 1127
LMArena Hard Prompts14471299
NYT Connections (extended)41.2%—

Math Hy3 leads

Hy3: 40.1 (#93), Qwen2.5 Plus 1127: 36.1 (#174)

Math benchmarks
BenchmarkHy3Qwen2.5 Plus 1127
LMArena Math14751298

Knowledge Hy3 leads

Hy3: 40.8 (#114), Qwen2.5 Plus 1127: 35.5 (#183)

Knowledge benchmarks
BenchmarkHy3Qwen2.5 Plus 1127
LMArena Expert14601289

Multilingual Hy3 leads

Hy3: 53.5 (#65), Qwen2.5 Plus 1127: 41.9 (#201)

Multilingual benchmarks
BenchmarkHy3Qwen2.5 Plus 1127
LMArena Non-English14261265
LMArena Chinese14931314
LMArena German14391231
LMArena Japanese13921207
LMArena Russian14321271
LMArena French1461—
LMArena Korean1395—
LMArena Spanish1456—

Instruction Following Hy3 leads

Hy3: 75.1 (#70), Qwen2.5 Plus 1127: 67.2 (#199)

Instruction Following benchmarks
BenchmarkHy3Qwen2.5 Plus 1127
LMArena Instruction Following14261275

Long Context Hy3 leads

Hy3: 44.1 (#75), Qwen2.5 Plus 1127: 39.2 (#184)

Long Context benchmarks
BenchmarkHy3Qwen2.5 Plus 1127
LMArena Longer Query14421292

Writing & Preference Hy3 leads

Hy3: 62.2 (#81), Qwen2.5 Plus 1127: 49.4 (#192)

Writing & Preference benchmarks
BenchmarkHy3Qwen2.5 Plus 1127
LMArena Text14391299
LMArena Creative Writing14021262
LMArena Multi-Turn14361299

Frequently asked questions

Is Hy3 better than Qwen2.5 Plus 1127?

Hy3 is the stronger model overall, scoring 44.2 to 38.8 on the Noometry Index.

Is Hy3 or Qwen2.5 Plus 1127 better for coding?

Hy3 scores higher on coding benchmarks: 46.8 versus 38.5 in the Noometry coding category.

How many benchmarks do Hy3 and Qwen2.5 Plus 1127 share?

14 benchmarks have published results for both models. Hy3 has 19 scored results on Noometry and Qwen2.5 Plus 1127 has 14.

Related comparisons

Go deeper