Model comparison

Llama 13b vs Mistral Nemo

Mistral Nemo is the stronger model overall, scoring 26.4 to 24.4 on the Noometry Index.

Last verified . 4 shared benchmarks.

Llama 13b Meta

24.4

Rank #348 Confirmed

Mistral Nemo Mistral AI

26.4

Rank #337 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Llama 13b scores higher in 1 category and Mistral Nemo in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Nemo leads 28.5 to 13.8.

Side by side

Llama 13b and Mistral Nemo specifications
Llama 13bMistral Nemo
ProviderMetaMistral AI
Noometry Index24.426.4
Released2023-02-242024-07-01
WeightsOpenOpen
Context window—128K
Max output—128K
Input $ / M tokens—$0.15
Output $ / M tokens—$0.15
Results tracked2110

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 13b: 21.4 (#337), Mistral Nemo: —

Coding benchmarks
BenchmarkLlama 13bMistral Nemo
LMArena Coding683—

Agentic & Tool Use Not comparable

Llama 13b: —, Mistral Nemo: 23.5 (#125)

Agentic & Tool Use benchmarks
BenchmarkLlama 13bMistral Nemo
Berkeley Function Calling Leaderboard—27.6%
BALROG—17.6%

Reasoning Mistral Nemo leads

Llama 13b: 14.0 (#329), Mistral Nemo: 20.7 (#232)

Reasoning benchmarks
BenchmarkLlama 13bMistral Nemo
Epoch Capabilities Index100.58118.68
PIQA80.1%83.5%
LMArena Hard Prompts728—
DTBench—48.6%
BIG-Bench Hard37.9%—
HellaSwag79.2%—
LAMBADA75.2%—
WinoGrande73%—

Math Llama 13b leads

Llama 13b: 26.7 (#256), Mistral Nemo: 25.5 (#268)

Math benchmarks
BenchmarkLlama 13bMistral Nemo
GSM8K20.6%84.2%
LMArena Math838—
MATH Level 5—10.8%

Knowledge Not comparable

Llama 13b: —, Mistral Nemo: 12.3 (#298)

Knowledge benchmarks
BenchmarkLlama 13bMistral Nemo
BoolQ78.7%82.5%
GPQA Diamond—29.9%
ARC (AI2) Challenge52.7%—
MMLU47.7%—
OpenBookQA56.4%—
TriviaQA77.9%—

Multimodal Not comparable

Llama 13b: —, Mistral Nemo: —

Multimodal benchmarks
BenchmarkLlama 13bMistral Nemo
ScienceQA43.3%—

Multilingual Not comparable

Llama 13b: 16.6 (#297), Mistral Nemo: —

Multilingual benchmarks
BenchmarkLlama 13bMistral Nemo
LMArena Non-English819—

Instruction Following Not comparable

Llama 13b: 36.7 (#305), Mistral Nemo: —

Instruction Following benchmarks
BenchmarkLlama 13bMistral Nemo
LMArena Instruction Following781—

Writing & Preference Mistral Nemo leads

Llama 13b: 13.8 (#312), Mistral Nemo: 28.5 (#296)

Writing & Preference benchmarks
BenchmarkLlama 13bMistral Nemo
LMArena Text834—
LMArena Creative Writing794—
EQ-Bench Creative Writing—881
LMArena Multi-Turn753—

Frequently asked questions

Is Llama 13b better than Mistral Nemo?

Mistral Nemo is the stronger model overall, scoring 26.4 to 24.4 on the Noometry Index.

How many benchmarks do Llama 13b and Mistral Nemo share?

4 benchmarks have published results for both models. Llama 13b has 21 scored results on Noometry and Mistral Nemo has 10.

Related comparisons

Go deeper