Model comparison

Llama 13b vs Qwen3.5 Max Preview

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 24.4 on the Noometry Index.

Last verified . 8 shared benchmarks.

Llama 13b Meta

24.4

Rank #348 Confirmed

Qwen3.5 Max Preview Alibaba (Qwen)

45.3

Rank #71 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Llama 13b scores higher in 0 categories and Qwen3.5 Max Preview in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen3.5 Max Preview leads 66.0 to 13.8.
  • Llama 13b has downloadable open weights; the other is API-only.

Side by side

Llama 13b and Qwen3.5 Max Preview specifications
Llama 13bQwen3.5 Max Preview
ProviderMetaAlibaba (Qwen)
Noometry Index24.445.3
Released2023-02-24—
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2117

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.5 Max Preview leads

Llama 13b: 21.4 (#337), Qwen3.5 Max Preview: 44.0 (#77)

Coding benchmarks
BenchmarkLlama 13bQwen3.5 Max Preview
LMArena Coding6831487

Reasoning Qwen3.5 Max Preview leads

Llama 13b: 14.0 (#329), Qwen3.5 Max Preview: 30.8 (#84)

Reasoning benchmarks
BenchmarkLlama 13bQwen3.5 Max Preview
LMArena Hard Prompts7281483
BIG-Bench Hard37.9%—
Epoch Capabilities Index100.58—
HellaSwag79.2%—
LAMBADA75.2%—
PIQA80.1%—
WinoGrande73%—

Math Qwen3.5 Max Preview leads

Llama 13b: 26.7 (#256), Qwen3.5 Max Preview: 40.1 (#94)

Math benchmarks
BenchmarkLlama 13bQwen3.5 Max Preview
LMArena Math8381474
GSM8K20.6%—

Knowledge Not comparable

Llama 13b: —, Qwen3.5 Max Preview: 41.8 (#107)

Knowledge benchmarks
BenchmarkLlama 13bQwen3.5 Max Preview
LMArena Expert—1489
ARC (AI2) Challenge52.7%—
BoolQ78.7%—
MMLU47.7%—
OpenBookQA56.4%—
TriviaQA77.9%—

Multimodal Not comparable

Llama 13b: —, Qwen3.5 Max Preview: —

Multimodal benchmarks
BenchmarkLlama 13bQwen3.5 Max Preview
ScienceQA43.3%—

Multilingual Qwen3.5 Max Preview leads

Llama 13b: 16.6 (#297), Qwen3.5 Max Preview: 56.2 (#22)

Multilingual benchmarks
BenchmarkLlama 13bQwen3.5 Max Preview
LMArena Non-English8191465
LMArena Chinese—1534
LMArena French—1484
LMArena German—1487
LMArena Japanese—1495
LMArena Korean—1438
LMArena Russian—1471
LMArena Spanish—1470

Instruction Following Qwen3.5 Max Preview leads

Llama 13b: 36.7 (#305), Qwen3.5 Max Preview: 77.0 (#31)

Instruction Following benchmarks
BenchmarkLlama 13bQwen3.5 Max Preview
LMArena Instruction Following7811467

Long Context Not comparable

Llama 13b: —, Qwen3.5 Max Preview: 45.2 (#45)

Long Context benchmarks
BenchmarkLlama 13bQwen3.5 Max Preview
LMArena Longer Query—1476

Writing & Preference Qwen3.5 Max Preview leads

Llama 13b: 13.8 (#312), Qwen3.5 Max Preview: 66.0 (#41)

Writing & Preference benchmarks
BenchmarkLlama 13bQwen3.5 Max Preview
LMArena Text8341470
LMArena Creative Writing7941464
LMArena Multi-Turn7531478

Frequently asked questions

Is Llama 13b better than Qwen3.5 Max Preview?

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 24.4 on the Noometry Index.

Is Llama 13b or Qwen3.5 Max Preview better for coding?

Qwen3.5 Max Preview scores higher on coding benchmarks: 44.0 versus 21.4 in the Noometry coding category.

How many benchmarks do Llama 13b and Qwen3.5 Max Preview share?

8 benchmarks have published results for both models. Llama 13b has 21 scored results on Noometry and Qwen3.5 Max Preview has 17.

Related comparisons

Go deeper