Model comparison

Llama 13b vs Qwen Max

Qwen Max is the stronger model overall, scoring 34.7 to 24.4 on the Noometry Index.

Last verified . 8 shared benchmarks.

Llama 13b Meta

24.4

Rank #348 Confirmed

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Llama 13b scores higher in 1 category and Qwen Max in 5 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen Max leads 47.8 to 13.8.
  • Llama 13b has downloadable open weights; the other is API-only.

Side by side

Llama 13b and Qwen Max specifications
Llama 13bQwen Max
ProviderMetaAlibaba (Qwen)
Noometry Index24.434.7
Released2023-02-242024-04-03
WeightsOpenProprietary
Context window—33K
Max output—8K
Input $ / M tokens—$1.60
Output $ / M tokens—$6.40
Results tracked2123

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen Max leads

Llama 13b: 21.4 (#337), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkLlama 13bQwen Max
LMArena Coding6831288
Aider Polyglot—21.8%

Reasoning Qwen Max leads

Llama 13b: 14.0 (#329), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkLlama 13bQwen Max
LMArena Hard Prompts7281269
BIG-Bench Hard37.9%—
Epoch Capabilities Index100.58—
HellaSwag79.2%—
LAMBADA75.2%—
PIQA80.1%—
WinoGrande73%—

Math Llama 13b leads

Llama 13b: 26.7 (#256), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkLlama 13bQwen Max
LMArena Math8381275
OTIS Mock AIME 2024-2025—16.1%
MATH Level 5—67.2%
FrontierMath (Feb 2025 set)—1%
GSM8K20.6%—

Knowledge Not comparable

Llama 13b: —, Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkLlama 13bQwen Max
GPQA Diamond—56.1%
LMArena Expert—1248
ARC (AI2) Challenge52.7%—
BoolQ78.7%—
MMLU47.7%—
OpenBookQA56.4%—
TriviaQA77.9%—

Multimodal Not comparable

Llama 13b: —, Qwen Max: —

Multimodal benchmarks
BenchmarkLlama 13bQwen Max
ScienceQA43.3%—

Multilingual Qwen Max leads

Llama 13b: 16.6 (#297), Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkLlama 13bQwen Max
LMArena Non-English8191263
LMArena Chinese—1254
LMArena French—1330
LMArena German—1254
LMArena Japanese—1205
LMArena Korean—1142
LMArena Russian—1274
LMArena Spanish—1290

Instruction Following Qwen Max leads

Llama 13b: 36.7 (#305), Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkLlama 13bQwen Max
LMArena Instruction Following7811262

Long Context Not comparable

Llama 13b: —, Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkLlama 13bQwen Max
Fiction.LiveBench—66.7%
LMArena Longer Query—1288

Writing & Preference Qwen Max leads

Llama 13b: 13.8 (#312), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkLlama 13bQwen Max
LMArena Text8341282
LMArena Creative Writing7941248
LMArena Multi-Turn7531277

Frequently asked questions

Is Llama 13b better than Qwen Max?

Qwen Max is the stronger model overall, scoring 34.7 to 24.4 on the Noometry Index.

Is Llama 13b or Qwen Max better for coding?

Qwen Max scores higher on coding benchmarks: 30.7 versus 21.4 in the Noometry coding category.

How many benchmarks do Llama 13b and Qwen Max share?

8 benchmarks have published results for both models. Llama 13b has 21 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper