Model comparison

Llama 13b vs Qwen3.6 Max Preview

Qwen3.6 Max Preview is the stronger model overall, scoring 51.5 to 24.4 on the Noometry Index.

Last verified . 9 shared benchmarks.

Llama 13b Meta

24.4

Rank #348 Confirmed

Qwen3.6 Max Preview Alibaba (Qwen)

51.5

Rank #43 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Llama 13b scores higher in 0 categories and Qwen3.6 Max Preview in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen3.6 Max Preview leads 63.8 to 13.8.
  • Llama 13b has downloadable open weights; the other is API-only.

Side by side

Llama 13b and Qwen3.6 Max Preview specifications
Llama 13bQwen3.6 Max Preview
ProviderMetaAlibaba (Qwen)
Noometry Index24.451.5
Released2023-02-242026-04-20
WeightsOpenProprietary
Context window—262K
Max output—66K
Input $ / M tokens—$1.30
Output $ / M tokens—$7.80
Results tracked2129

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.6 Max Preview leads

Llama 13b: 21.4 (#337), Qwen3.6 Max Preview: 48.7 (#54)

Coding benchmarks
BenchmarkLlama 13bQwen3.6 Max Preview
LMArena Coding6831471
SWE-bench Verified—76.7%
LMArena WebDev—1482

Agentic & Tool Use Not comparable

Llama 13b: —, Qwen3.6 Max Preview: —

Agentic & Tool Use benchmarks
BenchmarkLlama 13bQwen3.6 Max Preview
Vending-Bench 2—4,254

Reasoning Qwen3.6 Max Preview leads

Llama 13b: 14.0 (#329), Qwen3.6 Max Preview: 41.7 (#53)

Reasoning benchmarks
BenchmarkLlama 13bQwen3.6 Max Preview
LMArena Hard Prompts7281457
Epoch Capabilities Index100.58149.24
SimpleBench—63%
NYT Connections (extended)—74.1%
Chess Puzzles—20%
Mystery Game Puzzles—19%
DTBench—87.2%
LMCA—42.5%
BIG-Bench Hard37.9%—
HellaSwag79.2%—
LAMBADA75.2%—
PIQA80.1%—
WinoGrande73%—

Math Qwen3.6 Max Preview leads

Llama 13b: 26.7 (#256), Qwen3.6 Max Preview: 54.1 (#46)

Math benchmarks
BenchmarkLlama 13bQwen3.6 Max Preview
LMArena Math8381465
OTIS Mock AIME 2024-2025—91.1%
FrontierMath (Feb 2025 set)—23.1%
FrontierMath Tier 4 (v1)—4.2%
GSM8K20.6%—

Knowledge Not comparable

Llama 13b: —, Qwen3.6 Max Preview: 57.6 (#39)

Knowledge benchmarks
BenchmarkLlama 13bQwen3.6 Max Preview
GPQA Diamond—87.4%
SimpleQA Verified—52%
LMArena Expert—1478
ARC (AI2) Challenge52.7%—
BoolQ78.7%—
MMLU47.7%—
OpenBookQA56.4%—
TriviaQA77.9%—

Multimodal Not comparable

Llama 13b: —, Qwen3.6 Max Preview: —

Multimodal benchmarks
BenchmarkLlama 13bQwen3.6 Max Preview
ScienceQA43.3%—

Multilingual Qwen3.6 Max Preview leads

Llama 13b: 16.6 (#297), Qwen3.6 Max Preview: 54.2 (#48)

Multilingual benchmarks
BenchmarkLlama 13bQwen3.6 Max Preview
LMArena Non-English8191437
LMArena Chinese—1487
LMArena French—1449
LMArena Russian—1445
LMArena Spanish—1454

Instruction Following Qwen3.6 Max Preview leads

Llama 13b: 36.7 (#305), Qwen3.6 Max Preview: 75.7 (#55)

Instruction Following benchmarks
BenchmarkLlama 13bQwen3.6 Max Preview
LMArena Instruction Following7811438

Long Context Not comparable

Llama 13b: —, Qwen3.6 Max Preview: 44.6 (#61)

Long Context benchmarks
BenchmarkLlama 13bQwen3.6 Max Preview
LMArena Longer Query—1457

Writing & Preference Qwen3.6 Max Preview leads

Llama 13b: 13.8 (#312), Qwen3.6 Max Preview: 63.8 (#60)

Writing & Preference benchmarks
BenchmarkLlama 13bQwen3.6 Max Preview
LMArena Text8341447
LMArena Creative Writing7941435
LMArena Multi-Turn7531456

Frequently asked questions

Is Llama 13b better than Qwen3.6 Max Preview?

Qwen3.6 Max Preview is the stronger model overall, scoring 51.5 to 24.4 on the Noometry Index.

Is Llama 13b or Qwen3.6 Max Preview better for coding?

Qwen3.6 Max Preview scores higher on coding benchmarks: 48.7 versus 21.4 in the Noometry coding category.

How many benchmarks do Llama 13b and Qwen3.6 Max Preview share?

9 benchmarks have published results for both models. Llama 13b has 21 scored results on Noometry and Qwen3.6 Max Preview has 29.

Related comparisons

Go deeper