Model comparison
Falcon-180B vs Hunyuan Standard 2025 02 10
Hunyuan Standard 2025 02 10 is the stronger model overall, scoring 37.9 to 32.2 on the Noometry Index.
Last verified . 6 shared benchmarks.
Summary
- They share 6 benchmarks with published results for both. Falcon-180B scores higher in 0 categories and Hunyuan Standard 2025 02 10 in 4 categories; 4 gaps are clear of the uncertainty.
- The widest gap is in writing & preference, where Hunyuan Standard 2025 02 10 leads 47.2 to 29.1.
- Falcon-180B has downloadable open weights; the other is API-only.
Side by side
| Falcon-180B | Hunyuan Standard 2025 02 10 | |
|---|---|---|
| Provider | Technology Innovation Institute | Tencent |
| Noometry Index | 32.2 | 37.9 |
| Released | 2023-09-06 | — |
| Weights | Open | Proprietary |
| Context window | — | — |
| Max output | — | — |
| Input $ / M tokens | — | — |
| Output $ / M tokens | — | — |
| Results tracked | 16 | 12 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Not comparable
Falcon-180B: —, Hunyuan Standard 2025 02 10: 37.1 (#197)
| Benchmark | Falcon-180B | Hunyuan Standard 2025 02 10 |
|---|---|---|
| LMArena Coding | — | 1270 |
Reasoning Hunyuan Standard 2025 02 10 leads
Falcon-180B: 19.1 (#269), Hunyuan Standard 2025 02 10: 25.0 (#154)
| Benchmark | Falcon-180B | Hunyuan Standard 2025 02 10 |
|---|---|---|
| LMArena Hard Prompts | 1007 | 1264 |
| Epoch Capabilities Index | 112.13 | — |
| HellaSwag | 89% | — |
| LAMBADA | 79.8% | — |
| PIQA | 84.9% | — |
| WinoGrande | 87.1% | — |
Math Not comparable
Falcon-180B: —, Hunyuan Standard 2025 02 10: 35.6 (#179)
| Benchmark | Falcon-180B | Hunyuan Standard 2025 02 10 |
|---|---|---|
| LMArena Math | — | 1274 |
| GSM8K | 54.4% | — |
Knowledge Not comparable
Falcon-180B: —, Hunyuan Standard 2025 02 10: 34.2 (#197)
| Benchmark | Falcon-180B | Hunyuan Standard 2025 02 10 |
|---|---|---|
| LMArena Expert | — | 1248 |
| ARC (AI2) Challenge | 67.8% | — |
| BoolQ | 89% | — |
| MMLU | 70.6% | — |
| OpenBookQA | 64.2% | — |
Multilingual Hunyuan Standard 2025 02 10 leads
Falcon-180B: 25.2 (#286), Hunyuan Standard 2025 02 10: 41.7 (#203)
| Benchmark | Falcon-180B | Hunyuan Standard 2025 02 10 |
|---|---|---|
| LMArena Non-English | 1000 | 1262 |
| LMArena Chinese | — | 1319 |
| LMArena Russian | — | 1258 |
Instruction Following Hunyuan Standard 2025 02 10 leads
Falcon-180B: 53.4 (#286), Hunyuan Standard 2025 02 10: 65.5 (#219)
| Benchmark | Falcon-180B | Hunyuan Standard 2025 02 10 |
|---|---|---|
| LMArena Instruction Following | 1047 | 1245 |
Long Context Not comparable
Falcon-180B: —, Hunyuan Standard 2025 02 10: 39.5 (#173)
| Benchmark | Falcon-180B | Hunyuan Standard 2025 02 10 |
|---|---|---|
| LMArena Longer Query | — | 1301 |
Writing & Preference Hunyuan Standard 2025 02 10 leads
Falcon-180B: 29.1 (#295), Hunyuan Standard 2025 02 10: 47.2 (#214)
| Benchmark | Falcon-180B | Hunyuan Standard 2025 02 10 |
|---|---|---|
| LMArena Text | 1054 | 1274 |
| LMArena Creative Writing | 1089 | 1242 |
| LMArena Multi-Turn | 1013 | 1275 |
Frequently asked questions
Is Falcon-180B better than Hunyuan Standard 2025 02 10?
Hunyuan Standard 2025 02 10 is the stronger model overall, scoring 37.9 to 32.2 on the Noometry Index.
How many benchmarks do Falcon-180B and Hunyuan Standard 2025 02 10 share?
6 benchmarks have published results for both models. Falcon-180B has 16 scored results on Noometry and Hunyuan Standard 2025 02 10 has 12.