Model comparison
Llama 13b vs Mistral Nemo
Mistral Nemo is the stronger model overall, scoring 26.4 to 24.4 on the Noometry Index.
Last verified . 4 shared benchmarks.
Summary
- They share 4 benchmarks with published results for both. Llama 13b scores higher in 1 category and Mistral Nemo in 2 categories; 3 gaps are clear of the uncertainty.
- The widest gap is in writing & preference, where Mistral Nemo leads 28.5 to 13.8.
Side by side
| Llama 13b | Mistral Nemo | |
|---|---|---|
| Provider | Meta | Mistral AI |
| Noometry Index | 24.4 | 26.4 |
| Released | 2023-02-24 | 2024-07-01 |
| Weights | Open | Open |
| Context window | — | 128K |
| Max output | — | 128K |
| Input $ / M tokens | — | $0.15 |
| Output $ / M tokens | — | $0.15 |
| Results tracked | 21 | 10 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Not comparable
Llama 13b: 21.4 (#337), Mistral Nemo: —
| Benchmark | Llama 13b | Mistral Nemo |
|---|---|---|
| LMArena Coding | 683 | — |
Agentic & Tool Use Not comparable
Llama 13b: —, Mistral Nemo: 23.5 (#125)
| Benchmark | Llama 13b | Mistral Nemo |
|---|---|---|
| Berkeley Function Calling Leaderboard | — | 27.6% |
| BALROG | — | 17.6% |
Reasoning Mistral Nemo leads
Llama 13b: 14.0 (#329), Mistral Nemo: 20.7 (#232)
| Benchmark | Llama 13b | Mistral Nemo |
|---|---|---|
| Epoch Capabilities Index | 100.58 | 118.68 |
| PIQA | 80.1% | 83.5% |
| LMArena Hard Prompts | 728 | — |
| DTBench | — | 48.6% |
| BIG-Bench Hard | 37.9% | — |
| HellaSwag | 79.2% | — |
| LAMBADA | 75.2% | — |
| WinoGrande | 73% | — |
Math Llama 13b leads
Llama 13b: 26.7 (#256), Mistral Nemo: 25.5 (#268)
| Benchmark | Llama 13b | Mistral Nemo |
|---|---|---|
| GSM8K | 20.6% | 84.2% |
| LMArena Math | 838 | — |
| MATH Level 5 | — | 10.8% |
Knowledge Not comparable
Llama 13b: —, Mistral Nemo: 12.3 (#298)
| Benchmark | Llama 13b | Mistral Nemo |
|---|---|---|
| BoolQ | 78.7% | 82.5% |
| GPQA Diamond | — | 29.9% |
| ARC (AI2) Challenge | 52.7% | — |
| MMLU | 47.7% | — |
| OpenBookQA | 56.4% | — |
| TriviaQA | 77.9% | — |
Multimodal Not comparable
Llama 13b: —, Mistral Nemo: —
| Benchmark | Llama 13b | Mistral Nemo |
|---|---|---|
| ScienceQA | 43.3% | — |
Multilingual Not comparable
Llama 13b: 16.6 (#297), Mistral Nemo: —
| Benchmark | Llama 13b | Mistral Nemo |
|---|---|---|
| LMArena Non-English | 819 | — |
Instruction Following Not comparable
Llama 13b: 36.7 (#305), Mistral Nemo: —
| Benchmark | Llama 13b | Mistral Nemo |
|---|---|---|
| LMArena Instruction Following | 781 | — |
Writing & Preference Mistral Nemo leads
Llama 13b: 13.8 (#312), Mistral Nemo: 28.5 (#296)
| Benchmark | Llama 13b | Mistral Nemo |
|---|---|---|
| LMArena Text | 834 | — |
| LMArena Creative Writing | 794 | — |
| EQ-Bench Creative Writing | — | 881 |
| LMArena Multi-Turn | 753 | — |
Frequently asked questions
Is Llama 13b better than Mistral Nemo?
Mistral Nemo is the stronger model overall, scoring 26.4 to 24.4 on the Noometry Index.
How many benchmarks do Llama 13b and Mistral Nemo share?
4 benchmarks have published results for both models. Llama 13b has 21 scored results on Noometry and Mistral Nemo has 10.