Model comparison

Hy3 vs Qwen2-72B

Hy3 is the stronger model overall, scoring 44.2 to 30.0 on the Noometry Index.

Last verified . 17 shared benchmarks.

Hy3 Tencent

44.2

Rank #79 Confirmed

Qwen2-72B Alibaba (Qwen)

30.0

Rank #300 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Hy3 scores higher in 8 categories and Qwen2-72B in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Hy3 leads 62.2 to 40.8.

Side by side

Hy3 and Qwen2-72B specifications
Hy3Qwen2-72B
ProviderTencentAlibaba (Qwen)
Noometry Index44.230.0
Released2026-07-062024-06-07
WeightsOpenOpen
Context window262K—
Max output128K—
Input $ / M tokens$0.0825—
Output $ / M tokens$0.33—
Results tracked1926

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy3 leads

Hy3: 46.8 (#63), Qwen2-72B: 29.1 (#310)

Coding benchmarks
BenchmarkHy3Qwen2-72B
LMArena Coding14641196
LMArena WebDev1508—
WeirdML—11.3%
BigCodeBench Instruct—38.5%
BigCodeBench Complete—54%

Agentic & Tool Use Not comparable

Hy3: —, Qwen2-72B: 17.0 (#146)

Agentic & Tool Use benchmarks
BenchmarkHy3Qwen2-72B
TheAgentCompany—1.1%
METR Time Horizons—29.9%

Reasoning Hy3 leads

Hy3: 26.1 (#136), Qwen2-72B: 23.2 (#181)

Reasoning benchmarks
BenchmarkHy3Qwen2-72B
LMArena Hard Prompts14471191
NYT Connections (extended)41.2%—
Epoch Capabilities Index—125.28

Math Hy3 leads

Hy3: 40.1 (#93), Qwen2-72B: 30.2 (#236)

Math benchmarks
BenchmarkHy3Qwen2-72B
LMArena Math14751235
MATH Level 5—39.1%

Knowledge Hy3 leads

Hy3: 40.8 (#114), Qwen2-72B: 21.2 (#275)

Knowledge benchmarks
BenchmarkHy3Qwen2-72B
LMArena Expert14601171
GPQA Diamond—40.8%
MMLU—82.4%

Multilingual Hy3 leads

Hy3: 53.5 (#65), Qwen2-72B: 35.9 (#244)

Multilingual benchmarks
BenchmarkHy3Qwen2-72B
LMArena Non-English14261176
LMArena Chinese14931240
LMArena French14611170
LMArena German14391151
LMArena Japanese13921111
LMArena Korean13951083
LMArena Russian14321169
LMArena Spanish14561169

Instruction Following Hy3 leads

Hy3: 75.1 (#70), Qwen2-72B: 61.7 (#241)

Instruction Following benchmarks
BenchmarkHy3Qwen2-72B
LMArena Instruction Following14261181

Long Context Hy3 leads

Hy3: 44.1 (#75), Qwen2-72B: 36.1 (#235)

Long Context benchmarks
BenchmarkHy3Qwen2-72B
LMArena Longer Query14421192

Writing & Preference Hy3 leads

Hy3: 62.2 (#81), Qwen2-72B: 40.8 (#241)

Writing & Preference benchmarks
BenchmarkHy3Qwen2-72B
LMArena Text14391203
LMArena Creative Writing14021181
LMArena Multi-Turn14361196

Frequently asked questions

Is Hy3 better than Qwen2-72B?

Hy3 is the stronger model overall, scoring 44.2 to 30.0 on the Noometry Index.

Is Hy3 or Qwen2-72B better for coding?

Hy3 scores higher on coding benchmarks: 46.8 versus 29.1 in the Noometry coding category.

How many benchmarks do Hy3 and Qwen2-72B share?

17 benchmarks have published results for both models. Hy3 has 19 scored results on Noometry and Qwen2-72B has 26.

Related comparisons

Go deeper