Model comparison

Llama 13b vs Qwen Plus

Qwen Plus is the stronger model overall, scoring 37.1 to 24.4 on the Noometry Index.

Last verified . 8 shared benchmarks.

Llama 13b Meta

24.4

Rank #348 Confirmed

Qwen Plus Alibaba (Qwen)

37.1

Rank #210 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Llama 13b scores higher in 1 category and Qwen Plus in 5 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen Plus leads 52.2 to 13.8.
  • Llama 13b has downloadable open weights; the other is API-only.

Side by side

Llama 13b and Qwen Plus specifications
Llama 13bQwen Plus
ProviderMetaAlibaba (Qwen)
Noometry Index24.437.1
Released2023-02-242024-01-25
WeightsOpenProprietary
Context window—1M
Max output—33K
Input $ / M tokens—$0.40
Output $ / M tokens—$1.20
Results tracked2120

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen Plus leads

Llama 13b: 21.4 (#337), Qwen Plus: 38.9 (#167)

Coding benchmarks
BenchmarkLlama 13bQwen Plus
LMArena Coding6831328

Reasoning Qwen Plus leads

Llama 13b: 14.0 (#329), Qwen Plus: 28.4 (#107)

Reasoning benchmarks
BenchmarkLlama 13bQwen Plus
LMArena Hard Prompts7281317
Kagi LLM Benchmark—63.3%
DTBench—81.1%
LMCA—24%
BIG-Bench Hard37.9%—
Epoch Capabilities Index100.58—
HellaSwag79.2%—
LAMBADA75.2%—
PIQA80.1%—
WinoGrande73%—

Math Llama 13b leads

Llama 13b: 26.7 (#256), Qwen Plus: 23.3 (#271)

Math benchmarks
BenchmarkLlama 13bQwen Plus
LMArena Math8381326
OTIS Mock AIME 2024-2025—17.8%
MATH Level 5—65.3%
FrontierMath (Feb 2025 set)—1.7%
GSM8K20.6%—

Knowledge Not comparable

Llama 13b: —, Qwen Plus: 27.4 (#251)

Knowledge benchmarks
BenchmarkLlama 13bQwen Plus
GPQA Diamond—48.1%
LMArena Expert—1328
ARC (AI2) Challenge52.7%—
BoolQ78.7%—
MMLU47.7%—
OpenBookQA56.4%—
TriviaQA77.9%—

Multimodal Not comparable

Llama 13b: —, Qwen Plus: —

Multimodal benchmarks
BenchmarkLlama 13bQwen Plus
ScienceQA43.3%—

Multilingual Qwen Plus leads

Llama 13b: 16.6 (#297), Qwen Plus: 45.1 (#175)

Multilingual benchmarks
BenchmarkLlama 13bQwen Plus
LMArena Non-English8191310
LMArena Chinese—1347
LMArena Japanese—1251
LMArena Russian—1323

Instruction Following Qwen Plus leads

Llama 13b: 36.7 (#305), Qwen Plus: 68.8 (#181)

Instruction Following benchmarks
BenchmarkLlama 13bQwen Plus
LMArena Instruction Following7811303

Long Context Not comparable

Llama 13b: —, Qwen Plus: 40.3 (#158)

Long Context benchmarks
BenchmarkLlama 13bQwen Plus
LMArena Longer Query—1324

Writing & Preference Qwen Plus leads

Llama 13b: 13.8 (#312), Qwen Plus: 52.2 (#176)

Writing & Preference benchmarks
BenchmarkLlama 13bQwen Plus
LMArena Text8341326
LMArena Creative Writing7941293
LMArena Multi-Turn7531336

Frequently asked questions

Is Llama 13b better than Qwen Plus?

Qwen Plus is the stronger model overall, scoring 37.1 to 24.4 on the Noometry Index.

Is Llama 13b or Qwen Plus better for coding?

Qwen Plus scores higher on coding benchmarks: 38.9 versus 21.4 in the Noometry coding category.

How many benchmarks do Llama 13b and Qwen Plus share?

8 benchmarks have published results for both models. Llama 13b has 21 scored results on Noometry and Qwen Plus has 20.

Related comparisons

Go deeper