Model comparison
Llama 13b vs Qwen3.6 Flash
Qwen3.6 Flash is the stronger model overall, scoring 38.8 to 24.4 on the Noometry Index.
Last verified . 1 shared benchmarks.
Summary
- They share 1 benchmark with published results for both. Llama 13b scores higher in 0 categories and Qwen3.6 Flash in 2 categories; 2 gaps are clear of the uncertainty.
- The widest gap is in reasoning, where Qwen3.6 Flash leads 29.0 to 14.0.
- Llama 13b has downloadable open weights; the other is API-only.
Side by side
| Llama 13b | Qwen3.6 Flash | |
|---|---|---|
| Provider | Meta | Alibaba (Qwen) |
| Noometry Index | 24.4 | 38.8 |
| Released | 2023-02-24 | 2026-04-27 |
| Weights | Open | Proprietary |
| Context window | — | 1M |
| Max output | — | 66K |
| Input $ / M tokens | — | $0.19 |
| Output $ / M tokens | — | $1.13 |
| Results tracked | 21 | 13 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Not comparable
Llama 13b: 21.4 (#337), Qwen3.6 Flash: —
| Benchmark | Llama 13b | Qwen3.6 Flash |
|---|---|---|
| LMArena Coding | 683 | — |
| ALE-Bench | — | 326.4 |
Reasoning Qwen3.6 Flash leads
Llama 13b: 14.0 (#329), Qwen3.6 Flash: 29.0 (#96)
| Benchmark | Llama 13b | Qwen3.6 Flash |
|---|---|---|
| Epoch Capabilities Index | 100.58 | 143.26 |
| SimpleBench | — | 35.2% |
| Chess Puzzles | — | 20% |
| LMArena Hard Prompts | 728 | — |
| Mystery Game Puzzles | — | 18% |
| DTBench | — | 77.1% |
| LMCA | — | 31% |
| BIG-Bench Hard | 37.9% | — |
| HellaSwag | 79.2% | — |
| LAMBADA | 75.2% | — |
| PIQA | 80.1% | — |
| WinoGrande | 73% | — |
Math Qwen3.6 Flash leads
Llama 13b: 26.7 (#256), Qwen3.6 Flash: 39.0 (#117)
| Benchmark | Llama 13b | Qwen3.6 Flash |
|---|---|---|
| FrontierMath (Tiers 1-3) | — | 22.5% |
| OTIS Mock AIME 2024-2025 | — | 84.4% |
| LMArena Math | 838 | — |
| FrontierMath (Feb 2025 set) | — | 10.3% |
| FrontierMath Tier 4 (v1) | — | 0% |
| GSM8K | 20.6% | — |
Knowledge Not comparable
Llama 13b: —, Qwen3.6 Flash: 42.1 (#100)
| Benchmark | Llama 13b | Qwen3.6 Flash |
|---|---|---|
| GPQA Diamond | — | 83.3% |
| SimpleQA Verified | — | 15.9% |
| ARC (AI2) Challenge | 52.7% | — |
| BoolQ | 78.7% | — |
| MMLU | 47.7% | — |
| OpenBookQA | 56.4% | — |
| TriviaQA | 77.9% | — |
Multimodal Not comparable
Llama 13b: —, Qwen3.6 Flash: —
| Benchmark | Llama 13b | Qwen3.6 Flash |
|---|---|---|
| ScienceQA | 43.3% | — |
Multilingual Not comparable
Llama 13b: 16.6 (#297), Qwen3.6 Flash: —
| Benchmark | Llama 13b | Qwen3.6 Flash |
|---|---|---|
| LMArena Non-English | 819 | — |
Instruction Following Not comparable
Llama 13b: 36.7 (#305), Qwen3.6 Flash: —
| Benchmark | Llama 13b | Qwen3.6 Flash |
|---|---|---|
| LMArena Instruction Following | 781 | — |
Writing & Preference Not comparable
Llama 13b: 13.8 (#312), Qwen3.6 Flash: —
| Benchmark | Llama 13b | Qwen3.6 Flash |
|---|---|---|
| LMArena Text | 834 | — |
| LMArena Creative Writing | 794 | — |
| LMArena Multi-Turn | 753 | — |
Frequently asked questions
Is Llama 13b better than Qwen3.6 Flash?
Qwen3.6 Flash is the stronger model overall, scoring 38.8 to 24.4 on the Noometry Index.
How many benchmarks do Llama 13b and Qwen3.6 Flash share?
1 benchmark has published results for both models. Llama 13b has 21 scored results on Noometry and Qwen3.6 Flash has 13.