Model comparison

Llama 2-13B vs Qwen1.5 4b Chat

Llama 2-13B and Qwen1.5 4b Chat score almost the same on the Noometry Index (29.6 vs 28.8), so choose on price, context window or the category you care about most.

Last verified . 13 shared benchmarks.

Llama 2-13B Meta

29.6

Rank #309 Confirmed

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Llama 2-13B scores higher in 7 categories and Qwen1.5 4b Chat in 1 category; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Llama 2-13B leads 29.8 to 23.8.

Side by side

Llama 2-13B and Qwen1.5 4b Chat specifications
Llama 2-13BQwen1.5 4b Chat
ProviderMetaAlibaba (Qwen)
Noometry Index29.628.8
Released2023-07-18—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked3213

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 2-13B leads

Llama 2-13B: 30.9 (#291), Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkLlama 2-13BQwen1.5 4b Chat
LMArena Coding1062999

Reasoning Qwen1.5 4b Chat leads

Llama 2-13B: 12.8 (#337), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkLlama 2-13BQwen1.5 4b Chat
LMArena Hard Prompts1051976
Chess Puzzles0%—
DTBench42.2%—
BIG-Bench Hard58.2%—
Epoch Capabilities Index106.17—
HellaSwag80.7%—
LAMBADA76.5%—
PIQA80.8%—
WinoGrande72.8%—

Math Too close to call

Llama 2-13B: 31.1 (#229), Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkLlama 2-13BQwen1.5 4b Chat
LMArena Math10651026
GSM8K36.9%—

Knowledge Llama 2-13B leads

Llama 2-13B: 28.1 (#249), Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkLlama 2-13BQwen1.5 4b Chat
LMArena Expert1030980
ARC (AI2) Challenge60.3%—
BoolQ82.4%—
MMLU55.6%—
OpenBookQA57%—
TriviaQA79.6%—

Multimodal Not comparable

Llama 2-13B: —, Qwen1.5 4b Chat: —

Multimodal benchmarks
BenchmarkLlama 2-13BQwen1.5 4b Chat
ScienceQA55.8%—

Multilingual Llama 2-13B leads

Llama 2-13B: 26.5 (#279), Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkLlama 2-13BQwen1.5 4b Chat
LMArena Non-English1024979
LMArena Chinese10011024
LMArena German1009902
LMArena Russian1055952
LMArena French1044—
LMArena Japanese894—
LMArena Korean953—
LMArena Spanish1087—

Instruction Following Llama 2-13B leads

Llama 2-13B: 53.3 (#287), Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkLlama 2-13BQwen1.5 4b Chat
LMArena Instruction Following1045978

Long Context Llama 2-13B leads

Llama 2-13B: 32.3 (#269), Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkLlama 2-13BQwen1.5 4b Chat
LMArena Longer Query1064988

Writing & Preference Llama 2-13B leads

Llama 2-13B: 29.8 (#289), Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkLlama 2-13BQwen1.5 4b Chat
LMArena Text1084997
LMArena Creative Writing1047969
LMArena Multi-Turn1050977

Frequently asked questions

Is Llama 2-13B better than Qwen1.5 4b Chat?

Llama 2-13B and Qwen1.5 4b Chat score almost the same on the Noometry Index (29.6 vs 28.8), so choose on price, context window or the category you care about most.

Is Llama 2-13B or Qwen1.5 4b Chat better for coding?

Llama 2-13B scores higher on coding benchmarks: 30.9 versus 29.1 in the Noometry coding category.

How many benchmarks do Llama 2-13B and Qwen1.5 4b Chat share?

13 benchmarks have published results for both models. Llama 2-13B has 32 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper