Model comparison

Hy3 vs Llama2 70b Steerlm Chat

Hy3 is the stronger model overall, scoring 44.2 to 31.8 on the Noometry Index.

Last verified . 9 shared benchmarks.

Hy3 Tencent

44.2

Rank #79 Confirmed

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Hy3 scores higher in 7 categories and Llama2 70b Steerlm Chat in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Hy3 leads 62.2 to 31.6.

Side by side

Hy3 and Llama2 70b Steerlm Chat specifications
Hy3Llama2 70b Steerlm Chat
ProviderTencentNVIDIA
Noometry Index44.231.8
Released2026-07-06—
WeightsOpenOpen
Context window262K—
Max output128K—
Input $ / M tokens$0.0825—
Output $ / M tokens$0.33—
Results tracked199

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy3 leads

Hy3: 46.8 (#63), Llama2 70b Steerlm Chat: 29.9 (#300)

Coding benchmarks
BenchmarkHy3Llama2 70b Steerlm Chat
LMArena Coding14641025
LMArena WebDev1508—

Reasoning Hy3 leads

Hy3: 26.1 (#136), Llama2 70b Steerlm Chat: 20.0 (#246)

Reasoning benchmarks
BenchmarkHy3Llama2 70b Steerlm Chat
LMArena Hard Prompts14471047
NYT Connections (extended)41.2%—

Math Hy3 leads

Hy3: 40.1 (#93), Llama2 70b Steerlm Chat: 31.3 (#226)

Math benchmarks
BenchmarkHy3Llama2 70b Steerlm Chat
LMArena Math14751072

Knowledge Not comparable

Hy3: 40.8 (#114), Llama2 70b Steerlm Chat: —

Knowledge benchmarks
BenchmarkHy3Llama2 70b Steerlm Chat
LMArena Expert1460—

Multilingual Hy3 leads

Hy3: 53.5 (#65), Llama2 70b Steerlm Chat: 28.8 (#270)

Multilingual benchmarks
BenchmarkHy3Llama2 70b Steerlm Chat
LMArena Non-English14261063
LMArena Chinese1493—
LMArena French1461—
LMArena German1439—
LMArena Japanese1392—
LMArena Korean1395—
LMArena Russian1432—
LMArena Spanish1456—

Instruction Following Hy3 leads

Hy3: 75.1 (#70), Llama2 70b Steerlm Chat: 54.2 (#279)

Instruction Following benchmarks
BenchmarkHy3Llama2 70b Steerlm Chat
LMArena Instruction Following14261060

Long Context Hy3 leads

Hy3: 44.1 (#75), Llama2 70b Steerlm Chat: 30.4 (#288)

Long Context benchmarks
BenchmarkHy3Llama2 70b Steerlm Chat
LMArena Longer Query1442998

Writing & Preference Hy3 leads

Hy3: 62.2 (#81), Llama2 70b Steerlm Chat: 31.6 (#283)

Writing & Preference benchmarks
BenchmarkHy3Llama2 70b Steerlm Chat
LMArena Text14391098
LMArena Creative Writing14021091
LMArena Multi-Turn14361058

Frequently asked questions

Is Hy3 better than Llama2 70b Steerlm Chat?

Hy3 is the stronger model overall, scoring 44.2 to 31.8 on the Noometry Index.

Is Hy3 or Llama2 70b Steerlm Chat better for coding?

Hy3 scores higher on coding benchmarks: 46.8 versus 29.9 in the Noometry coding category.

How many benchmarks do Hy3 and Llama2 70b Steerlm Chat share?

9 benchmarks have published results for both models. Hy3 has 19 scored results on Noometry and Llama2 70b Steerlm Chat has 9.

Related comparisons

Go deeper