Model comparison

Llama 3-8B vs Mistral Nemo

Llama 3-8B and Mistral Nemo score almost the same on the Noometry Index (25.5 vs 26.4), so choose on price, context window or the category you care about most.

Last verified . 4 shared benchmarks.

Llama 3-8B Meta

25.5

Rank #344 Confirmed

Mistral Nemo Mistral AI

26.4

Rank #337 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Llama 3-8B scores higher in 1 category and Mistral Nemo in 3 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in math, where Mistral Nemo leads 25.5 to 8.8.

Side by side

Llama 3-8B and Mistral Nemo specifications
Llama 3-8BMistral Nemo
ProviderMetaMistral AI
Noometry Index25.526.4
Released2024-04-182024-07-01
WeightsOpenOpen
Context window—128K
Max output—128K
Input $ / M tokens—$0.15
Output $ / M tokens—$0.15
Results tracked3410

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 3-8B: 31.0 (#289), Mistral Nemo: —

Coding benchmarks
BenchmarkLlama 3-8BMistral Nemo
BigCodeBench Instruct31.9%—
LMArena Coding1152—
BigCodeBench Complete36.9%—
HumanEval+56.7%—
MBPP+54.8%—

Agentic & Tool Use Not comparable

Llama 3-8B: —, Mistral Nemo: 23.5 (#125)

Agentic & Tool Use benchmarks
BenchmarkLlama 3-8BMistral Nemo
Berkeley Function Calling Leaderboard—27.6%
BALROG—17.6%

Reasoning Mistral Nemo leads

Llama 3-8B: 14.3 (#326), Mistral Nemo: 20.7 (#232)

Reasoning benchmarks
BenchmarkLlama 3-8BMistral Nemo
DTBench43.9%48.6%
Epoch Capabilities Index116.45118.68
Chess Puzzles0%—
LMArena Hard Prompts1133—
Adversarial NLI57.3%—
ForecastBench58.6—
PIQA—83.5%
WinoGrande75.7%—

Math Mistral Nemo leads

Llama 3-8B: 8.8 (#323), Mistral Nemo: 25.5 (#268)

Math benchmarks
BenchmarkLlama 3-8BMistral Nemo
MATH Level 56.1%10.8%
OTIS Mock AIME 2024-20251.9%—
LMArena Math1151—
GSM8K—84.2%

Knowledge Mistral Nemo leads

Llama 3-8B: 7.8 (#308), Mistral Nemo: 12.3 (#298)

Knowledge benchmarks
BenchmarkLlama 3-8BMistral Nemo
GPQA Diamond26.1%29.9%
LMArena Expert1113—
ARC (AI2) Challenge82.8%—
BoolQ—82.5%
MMLU68.8%—
OpenBookQA82.6%—
TriviaQA67.7%—

Multilingual Not comparable

Llama 3-8B: 30.8 (#261), Mistral Nemo: —

Multilingual benchmarks
BenchmarkLlama 3-8BMistral Nemo
LMArena Non-English1098—
LMArena Chinese1076—
LMArena French1159—
LMArena German1104—
LMArena Japanese967—
LMArena Korean1004—
LMArena Russian1109—
LMArena Spanish1173—

Instruction Following Not comparable

Llama 3-8B: 58.4 (#260), Mistral Nemo: —

Instruction Following benchmarks
BenchmarkLlama 3-8BMistral Nemo
LMArena Instruction Following1127—

Long Context Not comparable

Llama 3-8B: 34.2 (#251), Mistral Nemo: —

Long Context benchmarks
BenchmarkLlama 3-8BMistral Nemo
LMArena Longer Query1128—

Writing & Preference Llama 3-8B leads

Llama 3-8B: 37.5 (#256), Mistral Nemo: 28.5 (#296)

Writing & Preference benchmarks
BenchmarkLlama 3-8BMistral Nemo
LMArena Text1166—
LMArena Creative Writing1150—
EQ-Bench Creative Writing—881
LMArena Multi-Turn1152—

Frequently asked questions

Is Llama 3-8B better than Mistral Nemo?

Llama 3-8B and Mistral Nemo score almost the same on the Noometry Index (25.5 vs 26.4), so choose on price, context window or the category you care about most.

How many benchmarks do Llama 3-8B and Mistral Nemo share?

4 benchmarks have published results for both models. Llama 3-8B has 34 scored results on Noometry and Mistral Nemo has 10.

Related comparisons

Go deeper