Model comparison
Claude 2 vs Llama 13b
Claude 2 and Llama 13b score almost the same on the Noometry Index (25.0 vs 24.4), so choose on price, context window or the category you care about most.
Last verified . 3 shared benchmarks.
Summary
- They share 3 benchmarks with published results for both. Claude 2 scores higher in 1 category and Llama 13b in 1 category; 2 gaps are clear of the uncertainty.
- The widest gap is in math, where Llama 13b leads 26.7 to 9.3.
- Llama 13b has downloadable open weights; the other is API-only.
Side by side
| Claude 2 | Llama 13b | |
|---|---|---|
| Provider | Anthropic | Meta |
| Noometry Index | 25.0 | 24.4 |
| Released | 2023-07-11 | 2023-02-24 |
| Weights | Proprietary | Open |
| Context window | — | — |
| Max output | — | — |
| Input $ / M tokens | — | — |
| Output $ / M tokens | — | — |
| Results tracked | 8 | 21 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Not comparable
Claude 2: —, Llama 13b: 21.4 (#337)
| Benchmark | Claude 2 | Llama 13b |
|---|---|---|
| LMArena Coding | — | 683 |
| HumanEval+ | 61.6% | — |
Reasoning Claude 2 leads
Claude 2: 21.7 (#216), Llama 13b: 14.0 (#329)
| Benchmark | Claude 2 | Llama 13b |
|---|---|---|
| Epoch Capabilities Index | 120.13 | 100.58 |
| LMArena Hard Prompts | — | 728 |
| DTBench | 51.9% | — |
| BIG-Bench Hard | — | 37.9% |
| HellaSwag | — | 79.2% |
| LAMBADA | — | 75.2% |
| PIQA | — | 80.1% |
| WinoGrande | — | 73% |
Math Llama 13b leads
Claude 2: 9.3 (#320), Llama 13b: 26.7 (#256)
| Benchmark | Claude 2 | Llama 13b |
|---|---|---|
| OTIS Mock AIME 2024-2025 | 2.5% | — |
| LMArena Math | — | 838 |
| MATH Level 5 | 11.7% | — |
| GSM8K | — | 20.6% |
Knowledge Not comparable
Claude 2: 16.9 (#287), Llama 13b: —
| Benchmark | Claude 2 | Llama 13b |
|---|---|---|
| MMLU | 78.5% | 47.7% |
| TriviaQA | 87.5% | 77.9% |
| GPQA Diamond | 34.7% | — |
| ARC (AI2) Challenge | — | 52.7% |
| BoolQ | — | 78.7% |
| OpenBookQA | — | 56.4% |
Multimodal Not comparable
Claude 2: —, Llama 13b: —
| Benchmark | Claude 2 | Llama 13b |
|---|---|---|
| ScienceQA | — | 43.3% |
Multilingual Not comparable
Claude 2: —, Llama 13b: 16.6 (#297)
| Benchmark | Claude 2 | Llama 13b |
|---|---|---|
| LMArena Non-English | — | 819 |
Instruction Following Not comparable
Claude 2: —, Llama 13b: 36.7 (#305)
| Benchmark | Claude 2 | Llama 13b |
|---|---|---|
| LMArena Instruction Following | — | 781 |
Writing & Preference Not comparable
Claude 2: —, Llama 13b: 13.8 (#312)
| Benchmark | Claude 2 | Llama 13b |
|---|---|---|
| LMArena Text | — | 834 |
| LMArena Creative Writing | — | 794 |
| LMArena Multi-Turn | — | 753 |
Frequently asked questions
Is Claude 2 better than Llama 13b?
Claude 2 and Llama 13b score almost the same on the Noometry Index (25.0 vs 24.4), so choose on price, context window or the category you care about most.
How many benchmarks do Claude 2 and Llama 13b share?
3 benchmarks have published results for both models. Claude 2 has 8 scored results on Noometry and Llama 13b has 21.