Model comparison

Hunyuan Turbos 20250226 vs Llama 13b

Hunyuan Turbos 20250226 is the stronger model overall, scoring 41.3 to 24.4 on the Noometry Index.

Last verified . 8 shared benchmarks.

Hunyuan Turbos 20250226 Tencent

41.3

Rank #139 Confirmed

Llama 13b Meta

24.4

Rank #348 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Hunyuan Turbos 20250226 scores higher in 6 categories and Llama 13b in 0 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Hunyuan Turbos 20250226 leads 57.4 to 13.8.
  • Llama 13b has downloadable open weights; the other is API-only.

Side by side

Hunyuan Turbos 20250226 and Llama 13b specifications
Hunyuan Turbos 20250226Llama 13b
ProviderTencentMeta
Noometry Index41.324.4
Released—2023-02-24
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1621

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Hunyuan Turbos 20250226 leads

Hunyuan Turbos 20250226: 40.0 (#152), Llama 13b: 21.4 (#337)

Coding benchmarks
BenchmarkHunyuan Turbos 20250226Llama 13b
LMArena Coding1361683

Reasoning Hunyuan Turbos 20250226 leads

Hunyuan Turbos 20250226: 27.8 (#113), Llama 13b: 14.0 (#329)

Reasoning benchmarks
BenchmarkHunyuan Turbos 20250226Llama 13b
LMArena Hard Prompts1374728
BIG-Bench Hard—37.9%
Epoch Capabilities Index—100.58
HellaSwag—79.2%
LAMBADA—75.2%
PIQA—80.1%
WinoGrande—73%

Math Hunyuan Turbos 20250226 leads

Hunyuan Turbos 20250226: 37.5 (#154), Llama 13b: 26.7 (#256)

Math benchmarks
BenchmarkHunyuan Turbos 20250226Llama 13b
LMArena Math1359838
GSM8K—20.6%

Knowledge Not comparable

Hunyuan Turbos 20250226: 37.0 (#161), Llama 13b: —

Knowledge benchmarks
BenchmarkHunyuan Turbos 20250226Llama 13b
LMArena Expert1339—
ARC (AI2) Challenge—52.7%
BoolQ—78.7%
MMLU—47.7%
OpenBookQA—56.4%
TriviaQA—77.9%

Multimodal Not comparable

Hunyuan Turbos 20250226: —, Llama 13b: —

Multimodal benchmarks
BenchmarkHunyuan Turbos 20250226Llama 13b
ScienceQA—43.3%

Multilingual Hunyuan Turbos 20250226 leads

Hunyuan Turbos 20250226: 48.9 (#136), Llama 13b: 16.6 (#297)

Multilingual benchmarks
BenchmarkHunyuan Turbos 20250226Llama 13b
LMArena Non-English1363819
LMArena Chinese1417—
LMArena French1391—
LMArena German1355—
LMArena Japanese1342—
LMArena Korean1351—
LMArena Russian1368—

Instruction Following Hunyuan Turbos 20250226 leads

Hunyuan Turbos 20250226: 71.0 (#158), Llama 13b: 36.7 (#305)

Instruction Following benchmarks
BenchmarkHunyuan Turbos 20250226Llama 13b
LMArena Instruction Following1344781

Long Context Not comparable

Hunyuan Turbos 20250226: 41.6 (#136), Llama 13b: —

Long Context benchmarks
BenchmarkHunyuan Turbos 20250226Llama 13b
LMArena Longer Query1366—

Writing & Preference Hunyuan Turbos 20250226 leads

Hunyuan Turbos 20250226: 57.4 (#128), Llama 13b: 13.8 (#312)

Writing & Preference benchmarks
BenchmarkHunyuan Turbos 20250226Llama 13b
LMArena Text1377834
LMArena Creative Writing1359794
LMArena Multi-Turn1387753

Frequently asked questions

Is Hunyuan Turbos 20250226 better than Llama 13b?

Hunyuan Turbos 20250226 is the stronger model overall, scoring 41.3 to 24.4 on the Noometry Index.

Is Hunyuan Turbos 20250226 or Llama 13b better for coding?

Hunyuan Turbos 20250226 scores higher on coding benchmarks: 40.0 versus 21.4 in the Noometry coding category.

How many benchmarks do Hunyuan Turbos 20250226 and Llama 13b share?

8 benchmarks have published results for both models. Hunyuan Turbos 20250226 has 16 scored results on Noometry and Llama 13b has 21.

Related comparisons

Go deeper