Model comparison

Llama 13b vs Qwen3.6 Flash

Qwen3.6 Flash is the stronger model overall, scoring 38.8 to 24.4 on the Noometry Index.

Last verified . 1 shared benchmarks.

Llama 13b Meta

24.4

Rank #348 Confirmed

Qwen3.6 Flash Alibaba (Qwen)

38.8

Rank #182 Confirmed

Summary

  • They share 1 benchmark with published results for both. Llama 13b scores higher in 0 categories and Qwen3.6 Flash in 2 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen3.6 Flash leads 29.0 to 14.0.
  • Llama 13b has downloadable open weights; the other is API-only.

Side by side

Llama 13b and Qwen3.6 Flash specifications
Llama 13bQwen3.6 Flash
ProviderMetaAlibaba (Qwen)
Noometry Index24.438.8
Released2023-02-242026-04-27
WeightsOpenProprietary
Context window—1M
Max output—66K
Input $ / M tokens—$0.19
Output $ / M tokens—$1.13
Results tracked2113

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 13b: 21.4 (#337), Qwen3.6 Flash: —

Coding benchmarks
BenchmarkLlama 13bQwen3.6 Flash
LMArena Coding683—
ALE-Bench—326.4

Reasoning Qwen3.6 Flash leads

Llama 13b: 14.0 (#329), Qwen3.6 Flash: 29.0 (#96)

Reasoning benchmarks
BenchmarkLlama 13bQwen3.6 Flash
Epoch Capabilities Index100.58143.26
SimpleBench—35.2%
Chess Puzzles—20%
LMArena Hard Prompts728—
Mystery Game Puzzles—18%
DTBench—77.1%
LMCA—31%
BIG-Bench Hard37.9%—
HellaSwag79.2%—
LAMBADA75.2%—
PIQA80.1%—
WinoGrande73%—

Math Qwen3.6 Flash leads

Llama 13b: 26.7 (#256), Qwen3.6 Flash: 39.0 (#117)

Math benchmarks
BenchmarkLlama 13bQwen3.6 Flash
FrontierMath (Tiers 1-3)—22.5%
OTIS Mock AIME 2024-2025—84.4%
LMArena Math838—
FrontierMath (Feb 2025 set)—10.3%
FrontierMath Tier 4 (v1)—0%
GSM8K20.6%—

Knowledge Not comparable

Llama 13b: —, Qwen3.6 Flash: 42.1 (#100)

Knowledge benchmarks
BenchmarkLlama 13bQwen3.6 Flash
GPQA Diamond—83.3%
SimpleQA Verified—15.9%
ARC (AI2) Challenge52.7%—
BoolQ78.7%—
MMLU47.7%—
OpenBookQA56.4%—
TriviaQA77.9%—

Multimodal Not comparable

Llama 13b: —, Qwen3.6 Flash: —

Multimodal benchmarks
BenchmarkLlama 13bQwen3.6 Flash
ScienceQA43.3%—

Multilingual Not comparable

Llama 13b: 16.6 (#297), Qwen3.6 Flash: —

Multilingual benchmarks
BenchmarkLlama 13bQwen3.6 Flash
LMArena Non-English819—

Instruction Following Not comparable

Llama 13b: 36.7 (#305), Qwen3.6 Flash: —

Instruction Following benchmarks
BenchmarkLlama 13bQwen3.6 Flash
LMArena Instruction Following781—

Writing & Preference Not comparable

Llama 13b: 13.8 (#312), Qwen3.6 Flash: —

Writing & Preference benchmarks
BenchmarkLlama 13bQwen3.6 Flash
LMArena Text834—
LMArena Creative Writing794—
LMArena Multi-Turn753—

Frequently asked questions

Is Llama 13b better than Qwen3.6 Flash?

Qwen3.6 Flash is the stronger model overall, scoring 38.8 to 24.4 on the Noometry Index.

How many benchmarks do Llama 13b and Qwen3.6 Flash share?

1 benchmark has published results for both models. Llama 13b has 21 scored results on Noometry and Qwen3.6 Flash has 13.

Related comparisons

Go deeper