Model comparison

Llama 3-8B vs Qwen2.5 72B Instruct

Qwen2.5 72B Instruct is the stronger model overall, scoring 31.9 to 25.5 on the Noometry Index.

Last verified . 29 shared benchmarks.

Llama 3-8B Meta

25.5

Rank #344 Confirmed

Qwen2.5 72B Instruct Alibaba (Qwen)

31.9

Rank #267 Confirmed

Summary

  • They share 29 benchmarks with published results for both. Llama 3-8B scores higher in 0 categories and Qwen2.5 72B Instruct in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen2.5 72B Instruct leads 27.0 to 7.8.
  • The biggest single-benchmark swing is MATH Level 5: 6.1% for Llama 3-8B and 63.2% for Qwen2.5 72B Instruct.

Side by side

Llama 3-8B and Qwen2.5 72B Instruct specifications
Llama 3-8BQwen2.5 72B Instruct
ProviderMetaAlibaba (Qwen)
Noometry Index25.531.9
Released2024-04-182024-09
WeightsOpenOpen
Context window—131K
Max output—8K
Input $ / M tokens—$1.40
Output $ / M tokens—$5.60
Results tracked3443

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 72B Instruct leads

Llama 3-8B: 31.0 (#289), Qwen2.5 72B Instruct: 33.2 (#260)

Coding benchmarks
BenchmarkLlama 3-8BQwen2.5 72B Instruct
BigCodeBench Instruct31.9%45.8%
LMArena Coding11521292
BigCodeBench Complete36.9%55.9%
WeirdML—16%
HumanEval+56.7%—
MBPP+54.8%—

Agentic & Tool Use Not comparable

Llama 3-8B: —, Qwen2.5 72B Instruct: 22.1 (#133)

Agentic & Tool Use benchmarks
BenchmarkLlama 3-8BQwen2.5 72B Instruct
TheAgentCompany—5.7%
BALROG—16.2%
METR Time Horizons—35.8%

Reasoning Qwen2.5 72B Instruct leads

Llama 3-8B: 14.3 (#326), Qwen2.5 72B Instruct: 22.3 (#199)

Reasoning benchmarks
BenchmarkLlama 3-8BQwen2.5 72B Instruct
LMArena Hard Prompts11331271
DTBench43.9%62.9%
Epoch Capabilities Index116.45129
ForecastBench58.657.5
WinoGrande75.7%82.3%
Chess Puzzles0%—
LMCA—13.4%
Adversarial NLI57.3%—
BIG-Bench Hard—79.8%
HellaSwag—84.8%
PIQA—82.6%

Math Qwen2.5 72B Instruct leads

Llama 3-8B: 8.8 (#323), Qwen2.5 72B Instruct: 19.3 (#287)

Math benchmarks
BenchmarkLlama 3-8BQwen2.5 72B Instruct
OTIS Mock AIME 2024-20251.9%8.1%
LMArena Math11511283
MATH Level 56.1%63.2%
Omni-MATH—33%

Knowledge Qwen2.5 72B Instruct leads

Llama 3-8B: 7.8 (#308), Qwen2.5 72B Instruct: 27.0 (#253)

Knowledge benchmarks
BenchmarkLlama 3-8BQwen2.5 72B Instruct
GPQA Diamond26.1%49.1%
LMArena Expert11131245
ARC (AI2) Challenge82.8%94.5%
MMLU68.8%85.3%
TriviaQA67.7%71.9%
MMLU-Pro—63.1%
Confabulations—19.1%
GPQA (HELM)—42.6%
OpenBookQA82.6%—

Multilingual Qwen2.5 72B Instruct leads

Llama 3-8B: 30.8 (#261), Qwen2.5 72B Instruct: 41.0 (#213)

Multilingual benchmarks
BenchmarkLlama 3-8BQwen2.5 72B Instruct
LMArena Non-English10981252
LMArena Chinese10761272
LMArena French11591280
LMArena German11041234
LMArena Japanese9671180
LMArena Korean10041188
LMArena Russian11091264
LMArena Spanish11731256

Instruction Following Qwen2.5 72B Instruct leads

Llama 3-8B: 58.4 (#260), Qwen2.5 72B Instruct: 65.5 (#221)

Instruction Following benchmarks
BenchmarkLlama 3-8BQwen2.5 72B Instruct
LMArena Instruction Following11271254
IFEval—80.6%

Long Context Qwen2.5 72B Instruct leads

Llama 3-8B: 34.2 (#251), Qwen2.5 72B Instruct: 38.9 (#188)

Long Context benchmarks
BenchmarkLlama 3-8BQwen2.5 72B Instruct
LMArena Longer Query11281282

Writing & Preference Qwen2.5 72B Instruct leads

Llama 3-8B: 37.5 (#256), Qwen2.5 72B Instruct: 46.7 (#215)

Writing & Preference benchmarks
BenchmarkLlama 3-8BQwen2.5 72B Instruct
LMArena Text11661269
LMArena Creative Writing11501221
LMArena Multi-Turn11521272
WildBench—80.2%

Frequently asked questions

Is Llama 3-8B better than Qwen2.5 72B Instruct?

Qwen2.5 72B Instruct is the stronger model overall, scoring 31.9 to 25.5 on the Noometry Index.

Is Llama 3-8B or Qwen2.5 72B Instruct better for coding?

Qwen2.5 72B Instruct scores higher on coding benchmarks: 33.2 versus 31.0 in the Noometry coding category.

How many benchmarks do Llama 3-8B and Qwen2.5 72B Instruct share?

29 benchmarks have published results for both models. Llama 3-8B has 34 scored results on Noometry and Qwen2.5 72B Instruct has 43.

Related comparisons

Go deeper