Model comparison

Llama 13b vs Ministral 3B

Ministral 3B is the stronger model overall, scoring 26.2 to 24.4 on the Noometry Index.

Last verified . 1 shared benchmarks.

Llama 13b Meta

24.4

Rank #348 Confirmed

Ministral 3B Mistral AI

26.2

Rank #338 Confirmed

Summary

  • They share 1 benchmark with published results for both. Llama 13b scores higher in 1 category and Ministral 3B in 1 category; one gap is clear of the uncertainty.
  • The widest gap is in reasoning, where Ministral 3B leads 18.4 to 14.0.

Side by side

Llama 13b and Ministral 3B specifications
Llama 13bMinistral 3B
ProviderMetaMistral AI
Noometry Index24.426.2
Released2023-02-242024-10-01
WeightsOpenOpen
Context window—131K
Max output—262K
Input $ / M tokens—$0.10
Output $ / M tokens—$0.10
Results tracked216

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 13b: 21.4 (#337), Ministral 3B: —

Coding benchmarks
BenchmarkLlama 13bMinistral 3B
LMArena Coding683—

Reasoning Ministral 3B leads

Llama 13b: 14.0 (#329), Ministral 3B: 18.4 (#282)

Reasoning benchmarks
BenchmarkLlama 13bMinistral 3B
Epoch Capabilities Index100.58118.1
LMArena Hard Prompts728—
DTBench—51.7%
LMCA—5.5%
BIG-Bench Hard37.9%—
HellaSwag79.2%—
LAMBADA75.2%—
PIQA80.1%—
WinoGrande73%—

Math Too close to call

Llama 13b: 26.7 (#256), Ministral 3B: 26.6 (#258)

Math benchmarks
BenchmarkLlama 13bMinistral 3B
LMArena Math838—
MATH Level 5—14.4%
GSM8K20.6%—

Knowledge Not comparable

Llama 13b: —, Ministral 3B: 10.4 (#302)

Knowledge benchmarks
BenchmarkLlama 13bMinistral 3B
GPQA Diamond—25.3%
Vectara Hallucination Rate—7.3%
ARC (AI2) Challenge52.7%—
BoolQ78.7%—
MMLU47.7%—
OpenBookQA56.4%—
TriviaQA77.9%—

Multimodal Not comparable

Llama 13b: —, Ministral 3B: —

Multimodal benchmarks
BenchmarkLlama 13bMinistral 3B
ScienceQA43.3%—

Multilingual Not comparable

Llama 13b: 16.6 (#297), Ministral 3B: —

Multilingual benchmarks
BenchmarkLlama 13bMinistral 3B
LMArena Non-English819—

Instruction Following Not comparable

Llama 13b: 36.7 (#305), Ministral 3B: —

Instruction Following benchmarks
BenchmarkLlama 13bMinistral 3B
LMArena Instruction Following781—

Writing & Preference Not comparable

Llama 13b: 13.8 (#312), Ministral 3B: —

Writing & Preference benchmarks
BenchmarkLlama 13bMinistral 3B
LMArena Text834—
LMArena Creative Writing794—
LMArena Multi-Turn753—

Frequently asked questions

Is Llama 13b better than Ministral 3B?

Ministral 3B is the stronger model overall, scoring 26.2 to 24.4 on the Noometry Index.

How many benchmarks do Llama 13b and Ministral 3B share?

1 benchmark has published results for both models. Llama 13b has 21 scored results on Noometry and Ministral 3B has 6.

Related comparisons

Go deeper