Model comparison

Llama 3.1-70B vs Qwen1.5 4b Chat

Llama 3.1-70B and Qwen1.5 4b Chat score almost the same on the Noometry Index (29.6 vs 28.8), so choose on price, context window or the category you care about most.

Last verified . 13 shared benchmarks.

Llama 3.1-70B Meta

29.6

Rank #308 Confirmed

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Llama 3.1-70B scores higher in 6 categories and Qwen1.5 4b Chat in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen1.5 4b Chat leads 30.4 to 13.5.

Side by side

Llama 3.1-70B and Qwen1.5 4b Chat specifications
Llama 3.1-70BQwen1.5 4b Chat
ProviderMetaAlibaba (Qwen)
Noometry Index29.628.8
Released2024-07-23—
WeightsOpenOpen
Context window128K—
Max output4K—
Input $ / M tokens$0.40—
Output $ / M tokens$0.40—
Results tracked3513

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.1-70B leads

Llama 3.1-70B: 30.3 (#296), Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkLlama 3.1-70BQwen1.5 4b Chat
LMArena Coding1260999
WeirdML9%—
BigCodeBench Instruct46.1%—
BigCodeBench Complete54.8%—

Agentic & Tool Use Not comparable

Llama 3.1-70B: 25.1 (#112), Qwen1.5 4b Chat: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.1-70BQwen1.5 4b Chat
TheAgentCompany6.9%—
BALROG27.9%—

Reasoning Llama 3.1-70B leads

Llama 3.1-70B: 21.6 (#220), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkLlama 3.1-70BQwen1.5 4b Chat
LMArena Hard Prompts1241976
DTBench60%—
LMCA14.8%—
Epoch Capabilities Index125.92—

Math Qwen1.5 4b Chat leads

Llama 3.1-70B: 13.5 (#304), Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkLlama 3.1-70BQwen1.5 4b Chat
LMArena Math12521026
OTIS Mock AIME 2024-20253.6%—
Omni-MATH21%—
MATH Level 536.7%—

Knowledge Qwen1.5 4b Chat leads

Llama 3.1-70B: 24.2 (#269), Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkLlama 3.1-70BQwen1.5 4b Chat
LMArena Expert1209980
GPQA Diamond44.2%—
MMLU-Pro65.3%—
GPQA (HELM)42.6%—
MMLU80.1%—

Multilingual Llama 3.1-70B leads

Llama 3.1-70B: 38.8 (#225), Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkLlama 3.1-70BQwen1.5 4b Chat
LMArena Non-English1219979
LMArena Chinese12151024
LMArena German1222902
LMArena Russian1234952
LMArena French1261—
LMArena Japanese1132—
LMArena Korean1140—
LMArena Spanish1253—

Instruction Following Llama 3.1-70B leads

Llama 3.1-70B: 65.3 (#223), Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkLlama 3.1-70BQwen1.5 4b Chat
LMArena Instruction Following1231978
IFEval82.1%—

Long Context Llama 3.1-70B leads

Llama 3.1-70B: 37.6 (#214), Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkLlama 3.1-70BQwen1.5 4b Chat
LMArena Longer Query1241988

Writing & Preference Llama 3.1-70B leads

Llama 3.1-70B: 35.4 (#267), Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkLlama 3.1-70BQwen1.5 4b Chat
LMArena Text1261997
LMArena Creative Writing1232969
LMArena Multi-Turn1256977
EQ-Bench Creative Writing784—
WildBench75.8%—

Frequently asked questions

Is Llama 3.1-70B better than Qwen1.5 4b Chat?

Llama 3.1-70B and Qwen1.5 4b Chat score almost the same on the Noometry Index (29.6 vs 28.8), so choose on price, context window or the category you care about most.

Is Llama 3.1-70B or Qwen1.5 4b Chat better for coding?

Llama 3.1-70B scores higher on coding benchmarks: 30.3 versus 29.1 in the Noometry coding category.

How many benchmarks do Llama 3.1-70B and Qwen1.5 4b Chat share?

13 benchmarks have published results for both models. Llama 3.1-70B has 35 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper