Model comparison

Llama 2-13B vs Qwen2.5 7B Instruct

Llama 2-13B and Qwen2.5 7B Instruct score almost the same on the Noometry Index (29.6 vs 29.0), so choose on price, context window or the category you care about most.

Last verified . 4 shared benchmarks.

Llama 2-13B Meta

29.6

Rank #309 Confirmed

Qwen2.5 7B Instruct Alibaba (Qwen)

29.0

Rank #320 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Llama 2-13B scores higher in 2 categories and Qwen2.5 7B Instruct in 4 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen2.5 7B Instruct leads 48.8 to 29.8.
  • The biggest single-benchmark swing is DTBench: 42.2% for Llama 2-13B and 47.7% for Qwen2.5 7B Instruct.

Side by side

Llama 2-13B and Qwen2.5 7B Instruct specifications
Llama 2-13BQwen2.5 7B Instruct
ProviderMetaAlibaba (Qwen)
Noometry Index29.629.0
Released2023-07-182024-09
WeightsOpenOpen
Context window—131K
Max output—8K
Input $ / M tokens—$0.17
Output $ / M tokens—$0.70
Results tracked3215

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 7B Instruct leads

Llama 2-13B: 30.9 (#291), Qwen2.5 7B Instruct: 36.5 (#208)

Coding benchmarks
BenchmarkLlama 2-13BQwen2.5 7B Instruct
BigCodeBench Instruct—37.6%
LMArena Coding1062—
BigCodeBench Complete—46.1%

Agentic & Tool Use Not comparable

Llama 2-13B: —, Qwen2.5 7B Instruct: 23.8 (#124)

Agentic & Tool Use benchmarks
BenchmarkLlama 2-13BQwen2.5 7B Instruct
BALROG—7.8%

Reasoning Qwen2.5 7B Instruct leads

Llama 2-13B: 12.8 (#337), Qwen2.5 7B Instruct: 14.8 (#322)

Reasoning benchmarks
BenchmarkLlama 2-13BQwen2.5 7B Instruct
Chess Puzzles0%0%
DTBench42.2%47.7%
Epoch Capabilities Index106.17118.51
LMArena Hard Prompts1051—
LMCA—6.4%
BIG-Bench Hard58.2%—
HellaSwag80.7%—
LAMBADA76.5%—
PIQA80.8%—
WinoGrande72.8%—

Math Llama 2-13B leads

Llama 2-13B: 31.1 (#229), Qwen2.5 7B Instruct: 12.6 (#306)

Math benchmarks
BenchmarkLlama 2-13BQwen2.5 7B Instruct
OTIS Mock AIME 2024-2025—2.5%
Omni-MATH—29.4%
LMArena Math1065—
GSM8K36.9%—

Knowledge Llama 2-13B leads

Llama 2-13B: 28.1 (#249), Qwen2.5 7B Instruct: 17.0 (#286)

Knowledge benchmarks
BenchmarkLlama 2-13BQwen2.5 7B Instruct
MMLU55.6%72.9%
GPQA Diamond—35.5%
MMLU-Pro—53.9%
GPQA (HELM)—34.1%
LMArena Expert1030—
ARC (AI2) Challenge60.3%—
BoolQ82.4%—
OpenBookQA57%—
TriviaQA79.6%—

Multimodal Not comparable

Llama 2-13B: —, Qwen2.5 7B Instruct: —

Multimodal benchmarks
BenchmarkLlama 2-13BQwen2.5 7B Instruct
ScienceQA55.8%—

Multilingual Not comparable

Llama 2-13B: 26.5 (#279), Qwen2.5 7B Instruct: —

Multilingual benchmarks
BenchmarkLlama 2-13BQwen2.5 7B Instruct
LMArena Non-English1024—
LMArena Chinese1001—
LMArena French1044—
LMArena German1009—
LMArena Japanese894—
LMArena Korean953—
LMArena Russian1055—
LMArena Spanish1087—

Instruction Following Qwen2.5 7B Instruct leads

Llama 2-13B: 53.3 (#287), Qwen2.5 7B Instruct: 63.2 (#231)

Instruction Following benchmarks
BenchmarkLlama 2-13BQwen2.5 7B Instruct
IFEval—74.1%
LMArena Instruction Following1045—

Long Context Not comparable

Llama 2-13B: 32.3 (#269), Qwen2.5 7B Instruct: —

Long Context benchmarks
BenchmarkLlama 2-13BQwen2.5 7B Instruct
LMArena Longer Query1064—

Writing & Preference Qwen2.5 7B Instruct leads

Llama 2-13B: 29.8 (#289), Qwen2.5 7B Instruct: 48.8 (#195)

Writing & Preference benchmarks
BenchmarkLlama 2-13BQwen2.5 7B Instruct
LMArena Text1084—
LMArena Creative Writing1047—
WildBench—73.1%
LMArena Multi-Turn1050—

Frequently asked questions

Is Llama 2-13B better than Qwen2.5 7B Instruct?

Llama 2-13B and Qwen2.5 7B Instruct score almost the same on the Noometry Index (29.6 vs 29.0), so choose on price, context window or the category you care about most.

Is Llama 2-13B or Qwen2.5 7B Instruct better for coding?

Qwen2.5 7B Instruct scores higher on coding benchmarks: 36.5 versus 30.9 in the Noometry coding category.

How many benchmarks do Llama 2-13B and Qwen2.5 7B Instruct share?

4 benchmarks have published results for both models. Llama 2-13B has 32 scored results on Noometry and Qwen2.5 7B Instruct has 15.

Related comparisons

Go deeper