Model comparison
Llama 3.1-405B vs Qwen2.5-VL 72B Instruct
Llama 3.1-405B and Qwen2.5-VL 72B Instruct score almost the same on the Noometry Index (30.7 vs 29.9), so choose on price, context window or the category you care about most.
Last verified . 1 shared benchmarks.
Summary
- They share 1 benchmark with published results for both. Llama 3.1-405B scores higher in 1 category and Qwen2.5-VL 72B Instruct in 1 category; 2 gaps are clear of the uncertainty.
- The widest gap is in reasoning, where Qwen2.5-VL 72B Instruct leads 20.7 to 16.8.
- The biggest single-benchmark swing is Kagi LLM Benchmark: 45% for Llama 3.1-405B and 36% for Qwen2.5-VL 72B Instruct.
Side by side
| Llama 3.1-405B | Qwen2.5-VL 72B Instruct | |
|---|---|---|
| Provider | Meta | Alibaba (Qwen) |
| Noometry Index | 30.7 | 29.9 |
| Released | 2024-07-23 | 2024-09 |
| Weights | Open | Open |
| Context window | — | 131K |
| Max output | — | 8K |
| Input $ / M tokens | — | $2.80 |
| Output $ / M tokens | — | $8.40 |
| Results tracked | 42 | 6 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Not comparable
Llama 3.1-405B: 33.1 (#262), Qwen2.5-VL 72B Instruct: —
| Benchmark | Llama 3.1-405B | Qwen2.5-VL 72B Instruct |
|---|---|---|
| WeirdML | 21.4% | — |
| LMArena Coding | 1291 | — |
Agentic & Tool Use Llama 3.1-405B leads
Llama 3.1-405B: 21.0 (#140), Qwen2.5-VL 72B Instruct: 18.6 (#144)
| Benchmark | Llama 3.1-405B | Qwen2.5-VL 72B Instruct |
|---|---|---|
| TheAgentCompany | 7.4% | — |
| Cybench | 7.5% | — |
| OSWorld | — | 5% |
Reasoning Qwen2.5-VL 72B Instruct leads
Llama 3.1-405B: 16.8 (#300), Qwen2.5-VL 72B Instruct: 20.7 (#233)
| Benchmark | Llama 3.1-405B | Qwen2.5-VL 72B Instruct |
|---|---|---|
| Kagi LLM Benchmark | 45% | 36% |
| SimpleBench | 23% | — |
| LMArena Hard Prompts | 1269 | — |
| DTBench | 61.4% | — |
| BIG-Bench Hard | 82.9% | — |
| Epoch Capabilities Index | 128.75 | — |
| ForecastBench | 59.9 | — |
| HellaSwag | 89.2% | — |
| PIQA | 85.9% | — |
| WinoGrande | 89.2% | — |
Math Not comparable
Llama 3.1-405B: 18.4 (#290), Qwen2.5-VL 72B Instruct: —
| Benchmark | Llama 3.1-405B | Qwen2.5-VL 72B Instruct |
|---|---|---|
| OTIS Mock AIME 2024-2025 | 9.7% | — |
| Omni-MATH | 24.9% | — |
| LMArena Math | 1281 | — |
| MATH Level 5 | 49.8% | — |
Knowledge Not comparable
Llama 3.1-405B: 30.4 (#227), Qwen2.5-VL 72B Instruct: —
| Benchmark | Llama 3.1-405B | Qwen2.5-VL 72B Instruct |
|---|---|---|
| GPQA Diamond | 50.9% | — |
| MMLU-Pro | 72.3% | — |
| Confabulations | 17.6% | — |
| GPQA (HELM) | 52.2% | — |
| LMArena Expert | 1243 | — |
| ARC (AI2) Challenge | 95.3% | — |
| MMLU | 84.5% | — |
| TriviaQA | 82.7% | — |
Multimodal Not comparable
Llama 3.1-405B: —, Qwen2.5-VL 72B Instruct: 33.5 (#97)
| Benchmark | Llama 3.1-405B | Qwen2.5-VL 72B Instruct |
|---|---|---|
| LMArena Vision | — | 1107 |
| Video-MME | — | 73.5% |
| GeoBench | — | 62% |
| SpatialViz-Bench | — | 33.3% |
Multilingual Not comparable
Llama 3.1-405B: 40.7 (#214), Qwen2.5-VL 72B Instruct: —
| Benchmark | Llama 3.1-405B | Qwen2.5-VL 72B Instruct |
|---|---|---|
| LMArena Non-English | 1248 | — |
| LMArena Chinese | 1242 | — |
| LMArena French | 1279 | — |
| LMArena German | 1252 | — |
| LMArena Japanese | 1208 | — |
| LMArena Korean | 1184 | — |
| LMArena Russian | 1265 | — |
| LMArena Spanish | 1260 | — |
Instruction Following Not comparable
Llama 3.1-405B: 65.9 (#214), Qwen2.5-VL 72B Instruct: —
| Benchmark | Llama 3.1-405B | Qwen2.5-VL 72B Instruct |
|---|---|---|
| IFEval | 81.1% | — |
| LMArena Instruction Following | 1259 | — |
Long Context Not comparable
Llama 3.1-405B: 38.4 (#197), Qwen2.5-VL 72B Instruct: —
| Benchmark | Llama 3.1-405B | Qwen2.5-VL 72B Instruct |
|---|---|---|
| LMArena Longer Query | 1266 | — |
Writing & Preference Not comparable
Llama 3.1-405B: 38.9 (#251), Qwen2.5-VL 72B Instruct: —
| Benchmark | Llama 3.1-405B | Qwen2.5-VL 72B Instruct |
|---|---|---|
| LMArena Text | 1284 | — |
| LMArena Creative Writing | 1262 | — |
| EQ-Bench Creative Writing | 870 | — |
| WildBench | 78.3% | — |
| LMArena Multi-Turn | 1297 | — |
Frequently asked questions
Is Llama 3.1-405B better than Qwen2.5-VL 72B Instruct?
Llama 3.1-405B and Qwen2.5-VL 72B Instruct score almost the same on the Noometry Index (30.7 vs 29.9), so choose on price, context window or the category you care about most.
How many benchmarks do Llama 3.1-405B and Qwen2.5-VL 72B Instruct share?
1 benchmark has published results for both models. Llama 3.1-405B has 42 scored results on Noometry and Qwen2.5-VL 72B Instruct has 6.