Model comparison

Llama 3.2 3B vs Mistral Large

Mistral Large is the stronger model overall, scoring 31.9 to 28.9 on the Noometry Index. Llama 3.2 3B costs 25× less per token, which makes it the better buy when Mistral Large's lead doesn't matter for your workload.

Last verified . 17 shared benchmarks.

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Llama 3.2 3B scores higher in 2 categories and Mistral Large in 7 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Large leads 40.7 to 24.7.
  • The biggest single-benchmark swing is Berkeley Function Calling Leaderboard: 21.9% for Llama 3.2 3B and 38.4% for Mistral Large.
  • Llama 3.2 3B is cheaper at $0.05 / $0.33 per million input/output tokens, against $2 / $6 for Mistral Large.

Side by side

Llama 3.2 3B and Mistral Large specifications
Llama 3.2 3BMistral Large
ProviderMetaMistral AI
Noometry Index28.931.9
Released2024-09-242024-02-26
WeightsOpenOpen
Context window131K131K
Max output118K16K
Input $ / M tokens$0.05$2
Output $ / M tokens$0.33$6
Results tracked1851

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Large leads

Llama 3.2 3B: 27.6 (#319), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkLlama 3.2 3BMistral Large
BigCodeBench Instruct23.4%30%
LMArena Coding10981277
BigCodeBench Complete28.3%38.3%
SciCode—36.2%
LiveBench Coding—47.1%
ALE-Bench—264.7
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Mistral Large leads

Llama 3.2 3B: 20.1 (#143), Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 3BMistral Large
Berkeley Function Calling Leaderboard21.9%38.4%
BALROG10.1%—

Reasoning Llama 3.2 3B leads

Llama 3.2 3B: 21.0 (#228), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkLlama 3.2 3BMistral Large
LMArena Hard Prompts10951257
SimpleBench—22.5%
CritPt—0%
LiveBench Reasoning—43.5%
DTBench—65.1%
LiveBench Data Analysis—50.1%
LMCA—16.7%
Epoch Capabilities Index—128.52
ForecastBench—57.1
LiveBench—48.4%

Math Llama 3.2 3B leads

Llama 3.2 3B: 32.4 (#214), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkLlama 3.2 3BMistral Large
LMArena Math11261262
OTIS Mock AIME 2024-2025—8.5%
Omni-MATH—28.1%
LiveBench Math—42.5%
MATH Level 5—50.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Too close to call

Llama 3.2 3B: 29.7 (#235), Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkLlama 3.2 3BMistral Large
LMArena Expert10901232
GPQA Diamond—51.3%
MMLU-Pro—59.9%
Confabulations—21.4%
Vectara Hallucination Rate—4.5%
GPQA (HELM)—43.5%
MMLU—80%

Multilingual Mistral Large leads

Llama 3.2 3B: 26.2 (#281), Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkLlama 3.2 3BMistral Large
LMArena Non-English10191237
LMArena Chinese10171240
LMArena German10561254
LMArena Russian9491257
LMArena French—1325
LMArena Japanese—1188
LMArena Korean—1202
LMArena Spanish—1268

Instruction Following Mistral Large leads

Llama 3.2 3B: 56.0 (#275), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkLlama 3.2 3BMistral Large
LMArena Instruction Following10891249
LiveBench Instruction Following—67.9%
IFEval—87.7%

Long Context Mistral Large leads

Llama 3.2 3B: 33.4 (#261), Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkLlama 3.2 3BMistral Large
LMArena Longer Query11001261

Writing & Preference Mistral Large leads

Llama 3.2 3B: 24.7 (#307), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkLlama 3.2 3BMistral Large
LMArena Text11101266
LMArena Creative Writing10941243
EQ-Bench Creative Writing595985
LMArena Multi-Turn11051260
Short-Story Creative Writing—69%
WildBench—80.1%
LiveBench Language—39.4%

Frequently asked questions

Is Llama 3.2 3B better than Mistral Large?

Mistral Large is the stronger model overall, scoring 31.9 to 28.9 on the Noometry Index. Llama 3.2 3B costs 25× less per token, which makes it the better buy when Mistral Large's lead doesn't matter for your workload.

Which is cheaper, Llama 3.2 3B or Mistral Large?

Llama 3.2 3B is cheaper. It lists at $0.05 per million input tokens and $0.33 per million output tokens; Mistral Large lists at $2 and $6.

Is Llama 3.2 3B or Mistral Large better for coding?

Mistral Large scores higher on coding benchmarks: 34.3 versus 27.6 in the Noometry coding category.

Which has the bigger context window?

Both accept 131K tokens.

How many benchmarks do Llama 3.2 3B and Mistral Large share?

17 benchmarks have published results for both models. Llama 3.2 3B has 18 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper