Model comparison

Llama 2-7B vs Qwen-14B

Qwen-14B is the stronger model overall, scoring 31.4 to 29.1 on the Noometry Index.

Last verified . 18 shared benchmarks.

Llama 2-7B Meta

29.1

Rank #317 Confirmed

Qwen-14B Alibaba (Qwen)

31.4

Rank #275 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Llama 2-7B scores higher in 1 category and Qwen-14B in 6 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen-14B leads 19.6 to 15.7.

Side by side

Llama 2-7B and Qwen-14B specifications
Llama 2-7BQwen-14B
ProviderMetaAlibaba (Qwen)
Noometry Index29.131.4
Released2023-07-182023-09-24
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2918

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen-14B leads

Llama 2-7B: 29.2 (#307), Qwen-14B: 31.2 (#288)

Coding benchmarks
BenchmarkLlama 2-7BQwen-14B
LMArena Coding10021071

Reasoning Qwen-14B leads

Llama 2-7B: 15.7 (#312), Qwen-14B: 19.6 (#257)

Reasoning benchmarks
BenchmarkLlama 2-7BQwen-14B
LMArena Hard Prompts10091027
BIG-Bench Hard39.2%55%
Epoch Capabilities Index99.06113.03
LAMBADA73.3%71.1%
PIQA78.8%79.9%
Chess Puzzles0%—
HellaSwag77.2%—
WinoGrande69.2%—

Math Too close to call

Llama 2-7B: 30.7 (#233), Qwen-14B: 31.2 (#227)

Math benchmarks
BenchmarkLlama 2-7BQwen-14B
LMArena Math10421068
GSM8K16.7%61.3%

Knowledge Not comparable

Llama 2-7B: 28.2 (#248), Qwen-14B: —

Knowledge benchmarks
BenchmarkLlama 2-7BQwen-14B
ARC (AI2) Challenge45.9%84.4%
BoolQ77.9%86.2%
MMLU45.8%66.3%
LMArena Expert1036—
OpenBookQA58.6%—
TriviaQA73.7%—

Multimodal Not comparable

Llama 2-7B: —, Qwen-14B: —

Multimodal benchmarks
BenchmarkLlama 2-7BQwen-14B
ScienceQA43.1%—

Multilingual Qwen-14B leads

Llama 2-7B: 23.8 (#293), Qwen-14B: 27.5 (#275)

Multilingual benchmarks
BenchmarkLlama 2-7BQwen-14B
LMArena Non-English9731041
LMArena Chinese9731077
LMArena French970—
LMArena German978—
LMArena Russian995—
LMArena Spanish1007—

Instruction Following Qwen-14B leads

Llama 2-7B: 50.8 (#298), Qwen-14B: 52.4 (#289)

Instruction Following benchmarks
BenchmarkLlama 2-7BQwen-14B
LMArena Instruction Following10061031

Long Context Too close to call

Llama 2-7B: 30.4 (#287), Qwen-14B: 31.3 (#280)

Long Context benchmarks
BenchmarkLlama 2-7BQwen-14B
LMArena Longer Query9991028

Writing & Preference Too close to call

Llama 2-7B: 28.0 (#298), Qwen-14B: 27.6 (#299)

Writing & Preference benchmarks
BenchmarkLlama 2-7BQwen-14B
LMArena Text10531051
LMArena Creative Writing10331028
LMArena Multi-Turn10291022

Frequently asked questions

Is Llama 2-7B better than Qwen-14B?

Qwen-14B is the stronger model overall, scoring 31.4 to 29.1 on the Noometry Index.

Is Llama 2-7B or Qwen-14B better for coding?

Qwen-14B scores higher on coding benchmarks: 31.2 versus 29.2 in the Noometry coding category.

How many benchmarks do Llama 2-7B and Qwen-14B share?

18 benchmarks have published results for both models. Llama 2-7B has 29 scored results on Noometry and Qwen-14B has 18.

Related comparisons

Go deeper