Model comparison

Hunyuan T1 20250711 vs Llama 13b

Hunyuan T1 20250711 is the stronger model overall, scoring 42.5 to 24.4 on the Noometry Index.

Last verified . 8 shared benchmarks.

Hunyuan T1 20250711 Tencent

42.5

Rank #114 Confirmed

Llama 13b Meta

24.4

Rank #348 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Hunyuan T1 20250711 scores higher in 6 categories and Llama 13b in 0 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Hunyuan T1 20250711 leads 59.5 to 13.8.
  • Llama 13b has downloadable open weights; the other is API-only.

Side by side

Hunyuan T1 20250711 and Llama 13b specifications
Hunyuan T1 20250711Llama 13b
ProviderTencentMeta
Noometry Index42.524.4
Released—2023-02-24
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1321

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hunyuan T1 20250711 leads

Hunyuan T1 20250711: 40.9 (#129), Llama 13b: 21.4 (#337)

Coding benchmarks
BenchmarkHunyuan T1 20250711Llama 13b
LMArena Coding1390683

Reasoning Hunyuan T1 20250711 leads

Hunyuan T1 20250711: 28.5 (#103), Llama 13b: 14.0 (#329)

Reasoning benchmarks
BenchmarkHunyuan T1 20250711Llama 13b
LMArena Hard Prompts1399728
BIG-Bench Hard—37.9%
Epoch Capabilities Index—100.58
HellaSwag—79.2%
LAMBADA—75.2%
PIQA—80.1%
WinoGrande—73%

Math Hunyuan T1 20250711 leads

Hunyuan T1 20250711: 38.7 (#130), Llama 13b: 26.7 (#256)

Math benchmarks
BenchmarkHunyuan T1 20250711Llama 13b
LMArena Math1414838
GSM8K—20.6%

Knowledge Not comparable

Hunyuan T1 20250711: 38.8 (#141), Llama 13b: —

Knowledge benchmarks
BenchmarkHunyuan T1 20250711Llama 13b
LMArena Expert1395—
ARC (AI2) Challenge—52.7%
BoolQ—78.7%
MMLU—47.7%
OpenBookQA—56.4%
TriviaQA—77.9%

Multimodal Not comparable

Hunyuan T1 20250711: —, Llama 13b: —

Multimodal benchmarks
BenchmarkHunyuan T1 20250711Llama 13b
ScienceQA—43.3%

Multilingual Hunyuan T1 20250711 leads

Hunyuan T1 20250711: 51.2 (#112), Llama 13b: 16.6 (#297)

Multilingual benchmarks
BenchmarkHunyuan T1 20250711Llama 13b
LMArena Non-English1395819
LMArena Chinese1425—
LMArena Korean1406—
LMArena Russian1385—

Instruction Following Hunyuan T1 20250711 leads

Hunyuan T1 20250711: 72.6 (#138), Llama 13b: 36.7 (#305)

Instruction Following benchmarks
BenchmarkHunyuan T1 20250711Llama 13b
LMArena Instruction Following1374781

Long Context Not comparable

Hunyuan T1 20250711: 42.2 (#128), Llama 13b: —

Long Context benchmarks
BenchmarkHunyuan T1 20250711Llama 13b
LMArena Longer Query1384—

Writing & Preference Hunyuan T1 20250711 leads

Hunyuan T1 20250711: 59.5 (#109), Llama 13b: 13.8 (#312)

Writing & Preference benchmarks
BenchmarkHunyuan T1 20250711Llama 13b
LMArena Text1401834
LMArena Creative Writing1392794
LMArena Multi-Turn1393753

Frequently asked questions

Is Hunyuan T1 20250711 better than Llama 13b?

Hunyuan T1 20250711 is the stronger model overall, scoring 42.5 to 24.4 on the Noometry Index.

Is Hunyuan T1 20250711 or Llama 13b better for coding?

Hunyuan T1 20250711 scores higher on coding benchmarks: 40.9 versus 21.4 in the Noometry coding category.

How many benchmarks do Hunyuan T1 20250711 and Llama 13b share?

8 benchmarks have published results for both models. Hunyuan T1 20250711 has 13 scored results on Noometry and Llama 13b has 21.

Related comparisons

Go deeper