# Claude 2.1 vs Llama 13b

> Claude 2.1 and Llama 13b score almost the same on the Noometry Index (25.2 vs 24.4), so choose on price, context window or the category you care about most.

- Canonical page: https://noometry.com/compare/claude-2-1-vs-llama-13b
- Last updated: 2026-10-10
- Shared benchmarks: 2

## Summary

- They share 2 benchmarks with published results for both. Claude 2.1 scores higher in 2 categories and Llama 13b in 1 category; 3 gaps are clear of the uncertainty.
- The widest gap is in math, where Llama 13b leads 26.7 to 10.2.
- Llama 13b has downloadable open weights; the other is API-only.

## Snapshot

| | Claude 2.1 | Llama 13b |
|---|---|---|
| Provider | Anthropic | Meta |
| Noometry Index | 25.2 | 24.4 |
| Rank | 345 | 348 |
| Context | — | — |
| Input $/M | — | — |
| Output $/M | — | — |
| Weights | Proprietary | Open |

## Coding

- Claude 2.1: 26.2 (#327)
- Llama 13b: 21.4 (#337)

| Benchmark | Claude 2.1 | Llama 13b |
|---|---|---|
| WeirdML | 7.1% | — |
| LMArena Coding | — | 683 |

## Reasoning

- Claude 2.1: 21.4 (#221)
- Llama 13b: 14.0 (#329)

| Benchmark | Claude 2.1 | Llama 13b |
|---|---|---|
| Epoch Capabilities Index | 119.27 | 100.58 |
| LMArena Hard Prompts | — | 728 |
| DTBench | 51% | — |
| BIG-Bench Hard | — | 37.9% |
| ForecastBench | 54.2 | — |
| HellaSwag | — | 79.2% |
| LAMBADA | — | 75.2% |
| PIQA | — | 80.1% |
| WinoGrande | — | 73% |

## Math

- Claude 2.1: 10.2 (#315)
- Llama 13b: 26.7 (#256)

| Benchmark | Claude 2.1 | Llama 13b |
|---|---|---|
| OTIS Mock AIME 2024-2025 | 1.9% | — |
| LMArena Math | — | 838 |
| GSM8K | — | 20.6% |

## Knowledge

- Claude 2.1: 15.4 (#292)
- Llama 13b: —

| Benchmark | Claude 2.1 | Llama 13b |
|---|---|---|
| MMLU | 73.5% | 47.7% |
| GPQA Diamond | 33% | — |
| ARC (AI2) Challenge | — | 52.7% |
| BoolQ | — | 78.7% |
| OpenBookQA | — | 56.4% |
| TriviaQA | — | 77.9% |

## Multimodal

- Claude 2.1: —
- Llama 13b: —

| Benchmark | Claude 2.1 | Llama 13b |
|---|---|---|
| ScienceQA | — | 43.3% |

## Multilingual

- Claude 2.1: —
- Llama 13b: 16.6 (#297)

| Benchmark | Claude 2.1 | Llama 13b |
|---|---|---|
| LMArena Non-English | — | 819 |

## Instruction Following

- Claude 2.1: —
- Llama 13b: 36.7 (#305)

| Benchmark | Claude 2.1 | Llama 13b |
|---|---|---|
| LMArena Instruction Following | — | 781 |

## Writing & Preference

- Claude 2.1: —
- Llama 13b: 13.8 (#312)

| Benchmark | Claude 2.1 | Llama 13b |
|---|---|---|
| LMArena Text | — | 834 |
| LMArena Creative Writing | — | 794 |
| LMArena Multi-Turn | — | 753 |

## FAQ

### Is Claude 2.1 better than Llama 13b?

Claude 2.1 and Llama 13b score almost the same on the Noometry Index (25.2 vs 24.4), so choose on price, context window or the category you care about most.

### Is Claude 2.1 or Llama 13b better for coding?

Claude 2.1 scores higher on coding benchmarks: 26.2 versus 21.4 in the Noometry coding category.

### How many benchmarks do Claude 2.1 and Llama 13b share?

2 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and Llama 13b has 21.
