Model comparison

Llama 3-70B vs Qwen1.5-32B

Qwen1.5-32B is the stronger model overall, scoring 30.5 to 28.8 on the Noometry Index.

Last verified . 21 shared benchmarks.

Llama 3-70B Meta

28.8

Rank #323 Confirmed

Qwen1.5-32B Alibaba (Qwen)

30.5

Rank #293 Confirmed

Summary

  • They share 21 benchmarks with published results for both. Llama 3-70B scores higher in 6 categories and Qwen1.5-32B in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen1.5-32B leads 33.0 to 12.8.
  • The biggest single-benchmark swing is BigCodeBench Complete: 54.5% for Llama 3-70B and 42% for Qwen1.5-32B.

Side by side

Llama 3-70B and Qwen1.5-32B specifications
Llama 3-70BQwen1.5-32B
ProviderMetaAlibaba (Qwen)
Noometry Index28.830.5
Released2024-04-182024-02-04
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked3121

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3-70B leads

Llama 3-70B: 35.8 (#218), Qwen1.5-32B: 31.7 (#282)

Coding benchmarks
BenchmarkLlama 3-70BQwen1.5-32B
BigCodeBench Instruct43.6%32.3%
LMArena Coding12061155
BigCodeBench Complete54.5%42%
HumanEval+72%—
MBPP+69%—

Agentic & Tool Use Not comparable

Llama 3-70B: 21.1 (#139), Qwen1.5-32B: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3-70BQwen1.5-32B
Cybench5%—

Reasoning Qwen1.5-32B leads

Llama 3-70B: 18.0 (#288), Qwen1.5-32B: 21.8 (#212)

Reasoning benchmarks
BenchmarkLlama 3-70BQwen1.5-32B
LMArena Hard Prompts11951130
Kagi LLM Benchmark35.1%—
DTBench54.2%—
Epoch Capabilities Index122.93—
ForecastBench57.1—
WinoGrande83.5%—

Math Qwen1.5-32B leads

Llama 3-70B: 12.8 (#305), Qwen1.5-32B: 33.0 (#207)

Math benchmarks
BenchmarkLlama 3-70BQwen1.5-32B
LMArena Math12181155
OTIS Mock AIME 2024-20254.3%—
MATH Level 522.6%—

Knowledge Llama 3-70B leads

Llama 3-70B: 20.8 (#277), Qwen1.5-32B: 13.5 (#296)

Knowledge benchmarks
BenchmarkLlama 3-70BQwen1.5-32B
GPQA Diamond40.6%30.7%
LMArena Expert11491126
MMLU79.3%74.4%

Multilingual Llama 3-70B leads

Llama 3-70B: 33.6 (#251), Qwen1.5-32B: 31.4 (#259)

Multilingual benchmarks
BenchmarkLlama 3-70BQwen1.5-32B
LMArena Non-English11421106
LMArena Chinese11141177
LMArena French12321101
LMArena German11691058
LMArena Japanese10171027
LMArena Korean10171008
LMArena Russian11591073
LMArena Spanish12411089

Instruction Following Llama 3-70B leads

Llama 3-70B: 62.5 (#238), Qwen1.5-32B: 57.7 (#265)

Instruction Following benchmarks
BenchmarkLlama 3-70BQwen1.5-32B
LMArena Instruction Following11941116

Long Context Too close to call

Llama 3-70B: 35.6 (#240), Qwen1.5-32B: 34.7 (#246)

Long Context benchmarks
BenchmarkLlama 3-70BQwen1.5-32B
LMArena Longer Query11741146

Writing & Preference Llama 3-70B leads

Llama 3-70B: 42.8 (#231), Qwen1.5-32B: 34.2 (#271)

Writing & Preference benchmarks
BenchmarkLlama 3-70BQwen1.5-32B
LMArena Text12211137
LMArena Creative Writing12101083
LMArena Multi-Turn12231140

Frequently asked questions

Is Llama 3-70B better than Qwen1.5-32B?

Qwen1.5-32B is the stronger model overall, scoring 30.5 to 28.8 on the Noometry Index.

Is Llama 3-70B or Qwen1.5-32B better for coding?

Llama 3-70B scores higher on coding benchmarks: 35.8 versus 31.7 in the Noometry coding category.

How many benchmarks do Llama 3-70B and Qwen1.5-32B share?

21 benchmarks have published results for both models. Llama 3-70B has 31 scored results on Noometry and Qwen1.5-32B has 21.

Related comparisons

Go deeper