Model comparison

Hy4 preview vs Llama 13b

Hy4 preview is the stronger model overall, scoring 45.3 to 24.4 on the Noometry Index.

Last verified . 0 shared benchmarks.

Hy4 preview Tencent

45.3

Rank #73 Reported

Llama 13b Meta

24.4

Rank #348 Confirmed

Summary

  • The widest gap is in coding, where Hy4 preview leads 51.6 to 21.4.

Side by side

Hy4 preview and Llama 13b specifications
Hy4 previewLlama 13b
ProviderTencentMeta
Noometry Index45.324.4
Released2026-08-282023-02-24
WeightsOpenOpen
Context window1.05M—
Max output64K—
Input $ / M tokens$0.83—
Output $ / M tokens$2.50—
Results tracked321

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hy4 preview leads

Hy4 preview: 51.6 (#38), Llama 13b: 21.4 (#337)

Coding benchmarks
BenchmarkHy4 previewLlama 13b
LMArena WebDev1632—
LMArena Coding—683

Reasoning Hy4 preview leads

Hy4 preview: 31.9 (#79), Llama 13b: 14.0 (#329)

Reasoning benchmarks
BenchmarkHy4 previewLlama 13b
NYT Connections (extended)68.2%—
LMArena Hard Prompts—728
BIG-Bench Hard—37.9%
Epoch Capabilities Index—100.58
HellaSwag—79.2%
LAMBADA—75.2%
PIQA—80.1%
WinoGrande—73%

Math Hy4 preview leads

Hy4 preview: 55.7 (#42), Llama 13b: 26.7 (#256)

Math benchmarks
BenchmarkHy4 previewLlama 13b
ProofBench75%—
LMArena Math—838
GSM8K—20.6%

Knowledge Not comparable

Hy4 preview: —, Llama 13b: —

Knowledge benchmarks
BenchmarkHy4 previewLlama 13b
ARC (AI2) Challenge—52.7%
BoolQ—78.7%
MMLU—47.7%
OpenBookQA—56.4%
TriviaQA—77.9%

Multimodal Not comparable

Hy4 preview: —, Llama 13b: —

Multimodal benchmarks
BenchmarkHy4 previewLlama 13b
ScienceQA—43.3%

Multilingual Not comparable

Hy4 preview: —, Llama 13b: 16.6 (#297)

Multilingual benchmarks
BenchmarkHy4 previewLlama 13b
LMArena Non-English—819

Instruction Following Not comparable

Hy4 preview: —, Llama 13b: 36.7 (#305)

Instruction Following benchmarks
BenchmarkHy4 previewLlama 13b
LMArena Instruction Following—781

Writing & Preference Not comparable

Hy4 preview: —, Llama 13b: 13.8 (#312)

Writing & Preference benchmarks
BenchmarkHy4 previewLlama 13b
LMArena Text—834
LMArena Creative Writing—794
LMArena Multi-Turn—753

Frequently asked questions

Is Hy4 preview better than Llama 13b?

Hy4 preview is the stronger model overall, scoring 45.3 to 24.4 on the Noometry Index.

Is Hy4 preview or Llama 13b better for coding?

Hy4 preview scores higher on coding benchmarks: 51.6 versus 21.4 in the Noometry coding category.

How many benchmarks do Hy4 preview and Llama 13b share?

0 benchmarks have published results for both models. Hy4 preview has 3 scored results on Noometry and Llama 13b has 21.

Related comparisons

Go deeper