Model comparison

Hy4 preview vs Llama2 70b Steerlm Chat

Hy4 preview is the stronger model overall, scoring 45.3 to 31.8 on the Noometry Index.

Last verified . 0 shared benchmarks.

Hy4 preview Tencent

45.3

Rank #73 Reported

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Summary

  • The widest gap is in math, where Hy4 preview leads 55.7 to 31.3.

Side by side

Hy4 preview and Llama2 70b Steerlm Chat specifications
Hy4 previewLlama2 70b Steerlm Chat
ProviderTencentNVIDIA
Noometry Index45.331.8
Released2026-08-28—
WeightsOpenOpen
Context window1.05M—
Max output64K—
Input $ / M tokens$0.83—
Output $ / M tokens$2.50—
Results tracked39

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy4 preview leads

Hy4 preview: 51.6 (#38), Llama2 70b Steerlm Chat: 29.9 (#300)

Coding benchmarks
BenchmarkHy4 previewLlama2 70b Steerlm Chat
LMArena WebDev1632—
LMArena Coding—1025

Reasoning Hy4 preview leads

Hy4 preview: 31.9 (#79), Llama2 70b Steerlm Chat: 20.0 (#246)

Reasoning benchmarks
BenchmarkHy4 previewLlama2 70b Steerlm Chat
NYT Connections (extended)68.2%—
LMArena Hard Prompts—1047

Math Hy4 preview leads

Hy4 preview: 55.7 (#42), Llama2 70b Steerlm Chat: 31.3 (#226)

Math benchmarks
BenchmarkHy4 previewLlama2 70b Steerlm Chat
ProofBench75%—
LMArena Math—1072

Multilingual Not comparable

Hy4 preview: —, Llama2 70b Steerlm Chat: 28.8 (#270)

Multilingual benchmarks
BenchmarkHy4 previewLlama2 70b Steerlm Chat
LMArena Non-English—1063

Instruction Following Not comparable

Hy4 preview: —, Llama2 70b Steerlm Chat: 54.2 (#279)

Instruction Following benchmarks
BenchmarkHy4 previewLlama2 70b Steerlm Chat
LMArena Instruction Following—1060

Long Context Not comparable

Hy4 preview: —, Llama2 70b Steerlm Chat: 30.4 (#288)

Long Context benchmarks
BenchmarkHy4 previewLlama2 70b Steerlm Chat
LMArena Longer Query—998

Writing & Preference Not comparable

Hy4 preview: —, Llama2 70b Steerlm Chat: 31.6 (#283)

Writing & Preference benchmarks
BenchmarkHy4 previewLlama2 70b Steerlm Chat
LMArena Text—1098
LMArena Creative Writing—1091
LMArena Multi-Turn—1058

Frequently asked questions

Is Hy4 preview better than Llama2 70b Steerlm Chat?

Hy4 preview is the stronger model overall, scoring 45.3 to 31.8 on the Noometry Index.

Is Hy4 preview or Llama2 70b Steerlm Chat better for coding?

Hy4 preview scores higher on coding benchmarks: 51.6 versus 29.9 in the Noometry coding category.

How many benchmarks do Hy4 preview and Llama2 70b Steerlm Chat share?

0 benchmarks have published results for both models. Hy4 preview has 3 scored results on Noometry and Llama2 70b Steerlm Chat has 9.

Related comparisons

Go deeper