Model comparison

Falcon-180B vs Hunyuan Standard 2025 02 10

Hunyuan Standard 2025 02 10 is the stronger model overall, scoring 37.9 to 32.2 on the Noometry Index.

Last verified . 6 shared benchmarks.

Summary

  • They share 6 benchmarks with published results for both. Falcon-180B scores higher in 0 categories and Hunyuan Standard 2025 02 10 in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Hunyuan Standard 2025 02 10 leads 47.2 to 29.1.
  • Falcon-180B has downloadable open weights; the other is API-only.

Side by side

Falcon-180B and Hunyuan Standard 2025 02 10 specifications
Falcon-180BHunyuan Standard 2025 02 10
ProviderTechnology Innovation InstituteTencent
Noometry Index32.237.9
Released2023-09-06—
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1612

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Falcon-180B: —, Hunyuan Standard 2025 02 10: 37.1 (#197)

Coding benchmarks
BenchmarkFalcon-180BHunyuan Standard 2025 02 10
LMArena Coding—1270

Reasoning Hunyuan Standard 2025 02 10 leads

Falcon-180B: 19.1 (#269), Hunyuan Standard 2025 02 10: 25.0 (#154)

Reasoning benchmarks
BenchmarkFalcon-180BHunyuan Standard 2025 02 10
LMArena Hard Prompts10071264
Epoch Capabilities Index112.13—
HellaSwag89%—
LAMBADA79.8%—
PIQA84.9%—
WinoGrande87.1%—

Math Not comparable

Falcon-180B: —, Hunyuan Standard 2025 02 10: 35.6 (#179)

Math benchmarks
BenchmarkFalcon-180BHunyuan Standard 2025 02 10
LMArena Math—1274
GSM8K54.4%—

Knowledge Not comparable

Falcon-180B: —, Hunyuan Standard 2025 02 10: 34.2 (#197)

Knowledge benchmarks
BenchmarkFalcon-180BHunyuan Standard 2025 02 10
LMArena Expert—1248
ARC (AI2) Challenge67.8%—
BoolQ89%—
MMLU70.6%—
OpenBookQA64.2%—

Multilingual Hunyuan Standard 2025 02 10 leads

Falcon-180B: 25.2 (#286), Hunyuan Standard 2025 02 10: 41.7 (#203)

Multilingual benchmarks
BenchmarkFalcon-180BHunyuan Standard 2025 02 10
LMArena Non-English10001262
LMArena Chinese—1319
LMArena Russian—1258

Instruction Following Hunyuan Standard 2025 02 10 leads

Falcon-180B: 53.4 (#286), Hunyuan Standard 2025 02 10: 65.5 (#219)

Instruction Following benchmarks
BenchmarkFalcon-180BHunyuan Standard 2025 02 10
LMArena Instruction Following10471245

Long Context Not comparable

Falcon-180B: —, Hunyuan Standard 2025 02 10: 39.5 (#173)

Long Context benchmarks
BenchmarkFalcon-180BHunyuan Standard 2025 02 10
LMArena Longer Query—1301

Writing & Preference Hunyuan Standard 2025 02 10 leads

Falcon-180B: 29.1 (#295), Hunyuan Standard 2025 02 10: 47.2 (#214)

Writing & Preference benchmarks
BenchmarkFalcon-180BHunyuan Standard 2025 02 10
LMArena Text10541274
LMArena Creative Writing10891242
LMArena Multi-Turn10131275

Frequently asked questions

Is Falcon-180B better than Hunyuan Standard 2025 02 10?

Hunyuan Standard 2025 02 10 is the stronger model overall, scoring 37.9 to 32.2 on the Noometry Index.

How many benchmarks do Falcon-180B and Hunyuan Standard 2025 02 10 share?

6 benchmarks have published results for both models. Falcon-180B has 16 scored results on Noometry and Hunyuan Standard 2025 02 10 has 12.

Related comparisons

Go deeper