Model comparison

Llama 13b vs Mistral Small 3

Mistral Small 3 is the stronger model overall, scoring 31.2 to 24.4 on the Noometry Index.

Last verified . 9 shared benchmarks.

Llama 13b Meta

24.4

Rank #348 Confirmed

Mistral Small 3 Mistral AI

31.2

Rank #278 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Llama 13b scores higher in 1 category and Mistral Small 3 in 5 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Mistral Small 3 leads 63.7 to 36.7.

Side by side

Llama 13b and Mistral Small 3 specifications
Llama 13bMistral Small 3
ProviderMetaMistral AI
Noometry Index24.431.2
Released2023-02-242025-01-30
WeightsOpenOpen
Context window—33K
Max output—16K
Input $ / M tokens—$0.05
Output $ / M tokens—$0.08
Results tracked2124

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Small 3 leads

Llama 13b: 21.4 (#337), Mistral Small 3: 36.5 (#207)

Coding benchmarks
BenchmarkLlama 13bMistral Small 3
LMArena Coding6831246
BigCodeBench Instruct—45.3%
BigCodeBench Complete—50.4%

Reasoning Mistral Small 3 leads

Llama 13b: 14.0 (#329), Mistral Small 3: 18.9 (#273)

Reasoning benchmarks
BenchmarkLlama 13bMistral Small 3
LMArena Hard Prompts7281233
Epoch Capabilities Index100.58127.07
Chess Puzzles—0%
BIG-Bench Hard37.9%—
HellaSwag79.2%—
LAMBADA75.2%—
PIQA80.1%—
WinoGrande73%—

Math Llama 13b leads

Llama 13b: 26.7 (#256), Mistral Small 3: 16.3 (#295)

Math benchmarks
BenchmarkLlama 13bMistral Small 3
LMArena Math8381240
OTIS Mock AIME 2024-2025—6.7%
GSM8K20.6%—

Knowledge Not comparable

Llama 13b: —, Mistral Small 3: 25.1 (#263)

Knowledge benchmarks
BenchmarkLlama 13bMistral Small 3
GPQA Diamond—47.3%
Confabulations—25.2%
LMArena Expert—1202
ARC (AI2) Challenge52.7%—
BoolQ78.7%—
MMLU47.7%—
OpenBookQA56.4%—
TriviaQA77.9%—

Multimodal Not comparable

Llama 13b: —, Mistral Small 3: —

Multimodal benchmarks
BenchmarkLlama 13bMistral Small 3
ScienceQA43.3%—

Multilingual Mistral Small 3 leads

Llama 13b: 16.6 (#297), Mistral Small 3: 37.3 (#236)

Multilingual benchmarks
BenchmarkLlama 13bMistral Small 3
LMArena Non-English8191198
LMArena Chinese—1204
LMArena French—1203
LMArena German—1211
LMArena Japanese—1111
LMArena Korean—1188
LMArena Russian—1216

Instruction Following Mistral Small 3 leads

Llama 13b: 36.7 (#305), Mistral Small 3: 63.7 (#229)

Instruction Following benchmarks
BenchmarkLlama 13bMistral Small 3
LMArena Instruction Following7811214

Long Context Not comparable

Llama 13b: —, Mistral Small 3: 37.8 (#211)

Long Context benchmarks
BenchmarkLlama 13bMistral Small 3
LMArena Longer Query—1246

Writing & Preference Mistral Small 3 leads

Llama 13b: 13.8 (#312), Mistral Small 3: 32.2 (#280)

Writing & Preference benchmarks
BenchmarkLlama 13bMistral Small 3
LMArena Text8341234
LMArena Creative Writing7941195
LMArena Multi-Turn7531217
EQ-Bench Creative Writing—707

Frequently asked questions

Is Llama 13b better than Mistral Small 3?

Mistral Small 3 is the stronger model overall, scoring 31.2 to 24.4 on the Noometry Index.

Is Llama 13b or Mistral Small 3 better for coding?

Mistral Small 3 scores higher on coding benchmarks: 36.5 versus 21.4 in the Noometry coding category.

How many benchmarks do Llama 13b and Mistral Small 3 share?

9 benchmarks have published results for both models. Llama 13b has 21 scored results on Noometry and Mistral Small 3 has 24.

Related comparisons

Go deeper