Model comparison

Llama 3.1-405B vs Qwen2.5-VL 72B Instruct

Llama 3.1-405B and Qwen2.5-VL 72B Instruct score almost the same on the Noometry Index (30.7 vs 29.9), so choose on price, context window or the category you care about most.

Last verified . 1 shared benchmarks.

Llama 3.1-405B Meta

30.7

Rank #288 Confirmed

Summary

  • They share 1 benchmark with published results for both. Llama 3.1-405B scores higher in 1 category and Qwen2.5-VL 72B Instruct in 1 category; 2 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen2.5-VL 72B Instruct leads 20.7 to 16.8.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 45% for Llama 3.1-405B and 36% for Qwen2.5-VL 72B Instruct.

Side by side

Llama 3.1-405B and Qwen2.5-VL 72B Instruct specifications
Llama 3.1-405BQwen2.5-VL 72B Instruct
ProviderMetaAlibaba (Qwen)
Noometry Index30.729.9
Released2024-07-232024-09
WeightsOpenOpen
Context window—131K
Max output—8K
Input $ / M tokens—$2.80
Output $ / M tokens—$8.40
Results tracked426

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 3.1-405B: 33.1 (#262), Qwen2.5-VL 72B Instruct: —

Coding benchmarks
BenchmarkLlama 3.1-405BQwen2.5-VL 72B Instruct
WeirdML21.4%—
LMArena Coding1291—

Agentic & Tool Use Llama 3.1-405B leads

Llama 3.1-405B: 21.0 (#140), Qwen2.5-VL 72B Instruct: 18.6 (#144)

Agentic & Tool Use benchmarks
BenchmarkLlama 3.1-405BQwen2.5-VL 72B Instruct
TheAgentCompany7.4%—
Cybench7.5%—
OSWorld—5%

Reasoning Qwen2.5-VL 72B Instruct leads

Llama 3.1-405B: 16.8 (#300), Qwen2.5-VL 72B Instruct: 20.7 (#233)

Reasoning benchmarks
BenchmarkLlama 3.1-405BQwen2.5-VL 72B Instruct
Kagi LLM Benchmark45%36%
SimpleBench23%—
LMArena Hard Prompts1269—
DTBench61.4%—
BIG-Bench Hard82.9%—
Epoch Capabilities Index128.75—
ForecastBench59.9—
HellaSwag89.2%—
PIQA85.9%—
WinoGrande89.2%—

Math Not comparable

Llama 3.1-405B: 18.4 (#290), Qwen2.5-VL 72B Instruct: —

Math benchmarks
BenchmarkLlama 3.1-405BQwen2.5-VL 72B Instruct
OTIS Mock AIME 2024-20259.7%—
Omni-MATH24.9%—
LMArena Math1281—
MATH Level 549.8%—

Knowledge Not comparable

Llama 3.1-405B: 30.4 (#227), Qwen2.5-VL 72B Instruct: —

Knowledge benchmarks
BenchmarkLlama 3.1-405BQwen2.5-VL 72B Instruct
GPQA Diamond50.9%—
MMLU-Pro72.3%—
Confabulations17.6%—
GPQA (HELM)52.2%—
LMArena Expert1243—
ARC (AI2) Challenge95.3%—
MMLU84.5%—
TriviaQA82.7%—

Multimodal Not comparable

Llama 3.1-405B: —, Qwen2.5-VL 72B Instruct: 33.5 (#97)

Multimodal benchmarks
BenchmarkLlama 3.1-405BQwen2.5-VL 72B Instruct
LMArena Vision—1107
Video-MME—73.5%
GeoBench—62%
SpatialViz-Bench—33.3%

Multilingual Not comparable

Llama 3.1-405B: 40.7 (#214), Qwen2.5-VL 72B Instruct: —

Multilingual benchmarks
BenchmarkLlama 3.1-405BQwen2.5-VL 72B Instruct
LMArena Non-English1248—
LMArena Chinese1242—
LMArena French1279—
LMArena German1252—
LMArena Japanese1208—
LMArena Korean1184—
LMArena Russian1265—
LMArena Spanish1260—

Instruction Following Not comparable

Llama 3.1-405B: 65.9 (#214), Qwen2.5-VL 72B Instruct: —

Instruction Following benchmarks
BenchmarkLlama 3.1-405BQwen2.5-VL 72B Instruct
IFEval81.1%—
LMArena Instruction Following1259—

Long Context Not comparable

Llama 3.1-405B: 38.4 (#197), Qwen2.5-VL 72B Instruct: —

Long Context benchmarks
BenchmarkLlama 3.1-405BQwen2.5-VL 72B Instruct
LMArena Longer Query1266—

Writing & Preference Not comparable

Llama 3.1-405B: 38.9 (#251), Qwen2.5-VL 72B Instruct: —

Writing & Preference benchmarks
BenchmarkLlama 3.1-405BQwen2.5-VL 72B Instruct
LMArena Text1284—
LMArena Creative Writing1262—
EQ-Bench Creative Writing870—
WildBench78.3%—
LMArena Multi-Turn1297—

Frequently asked questions

Is Llama 3.1-405B better than Qwen2.5-VL 72B Instruct?

Llama 3.1-405B and Qwen2.5-VL 72B Instruct score almost the same on the Noometry Index (30.7 vs 29.9), so choose on price, context window or the category you care about most.

How many benchmarks do Llama 3.1-405B and Qwen2.5-VL 72B Instruct share?

1 benchmark has published results for both models. Llama 3.1-405B has 42 scored results on Noometry and Qwen2.5-VL 72B Instruct has 6.

Related comparisons

Go deeper