Model comparison

Llama 3.2 1B vs Mistral Large

Mistral Large is the stronger model overall, scoring 31.9 to 20.1 on the Noometry Index. Llama 3.2 1B costs 43× less per token, which makes it the better buy when Mistral Large's lead doesn't matter for your workload.

Last verified . 20 shared benchmarks.

Llama 3.2 1B Meta

20.1

Rank #354 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 20 benchmarks with published results for both. Llama 3.2 1B scores higher in 1 category and Mistral Large in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Mistral Large leads 30.1 to 7.2.
  • The biggest single-benchmark swing is Berkeley Function Calling Leaderboard: 10.8% for Llama 3.2 1B and 38.4% for Mistral Large.
  • Llama 3.2 1B is cheaper at $0.027 / $0.20 per million input/output tokens, against $2 / $6 for Mistral Large.
  • Mistral Large accepts more context: 131K tokens versus 60K.

Side by side

Llama 3.2 1B and Mistral Large specifications
Llama 3.2 1BMistral Large
ProviderMetaMistral AI
Noometry Index20.131.9
Released2024-09-242024-02-26
WeightsOpenOpen
Context window60K131K
Max output54K16K
Input $ / M tokens$0.027$2
Output $ / M tokens$0.20$6
Results tracked2251

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Large leads

Llama 3.2 1B: 21.1 (#338), Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkLlama 3.2 1BMistral Large
BigCodeBench Instruct8.2%30%
LMArena Coding10701277
BigCodeBench Complete11.3%38.3%
SciCode—36.2%
LiveBench Coding—47.1%
ALE-Bench—264.7
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Mistral Large leads

Llama 3.2 1B: 14.6 (#150), Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 1BMistral Large
Berkeley Function Calling Leaderboard10.8%38.4%
BALROG6.6%—

Reasoning Too close to call

Llama 3.2 1B: 16.2 (#308), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkLlama 3.2 1BMistral Large
LMArena Hard Prompts10441257
Epoch Capabilities Index101.99128.52
SimpleBench—22.5%
CritPt—0%
Chess Puzzles0%—
LiveBench Reasoning—43.5%
DTBench—65.1%
LiveBench Data Analysis—50.1%
LMCA—16.7%
ForecastBench—57.1
LiveBench—48.4%

Math Mistral Large leads

Llama 3.2 1B: 10.4 (#313), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkLlama 3.2 1BMistral Large
OTIS Mock AIME 2024-20250.6%8.5%
LMArena Math10861262
Omni-MATH—28.1%
LiveBench Math—42.5%
MATH Level 5—50.3%
FrontierMath (Feb 2025 set)—0.3%

Knowledge Mistral Large leads

Llama 3.2 1B: 7.2 (#312), Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkLlama 3.2 1BMistral Large
GPQA Diamond23.9%51.3%
LMArena Expert10071232
MMLU-Pro—59.9%
Confabulations—21.4%
Vectara Hallucination Rate—4.5%
GPQA (HELM)—43.5%
MMLU—80%

Multilingual Mistral Large leads

Llama 3.2 1B: 23.8 (#292), Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkLlama 3.2 1BMistral Large
LMArena Non-English9731237
LMArena Chinese9591240
LMArena German10141254
LMArena Russian9411257
LMArena French—1325
LMArena Japanese—1188
LMArena Korean—1202
LMArena Spanish—1268

Instruction Following Mistral Large leads

Llama 3.2 1B: 52.4 (#290), Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkLlama 3.2 1BMistral Large
LMArena Instruction Following10311249
LiveBench Instruction Following—67.9%
IFEval—87.7%

Long Context Mistral Large leads

Llama 3.2 1B: 31.9 (#274), Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkLlama 3.2 1BMistral Large
LMArena Longer Query10501261

Writing & Preference Mistral Large leads

Llama 3.2 1B: 21.3 (#310), Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkLlama 3.2 1BMistral Large
LMArena Text10551266
LMArena Creative Writing10331243
EQ-Bench Creative Writing200985
LMArena Multi-Turn10301260
Short-Story Creative Writing—69%
WildBench—80.1%
LiveBench Language—39.4%

Frequently asked questions

Is Llama 3.2 1B better than Mistral Large?

Mistral Large is the stronger model overall, scoring 31.9 to 20.1 on the Noometry Index. Llama 3.2 1B costs 43× less per token, which makes it the better buy when Mistral Large's lead doesn't matter for your workload.

Which is cheaper, Llama 3.2 1B or Mistral Large?

Llama 3.2 1B is cheaper. It lists at $0.027 per million input tokens and $0.20 per million output tokens; Mistral Large lists at $2 and $6.

Is Llama 3.2 1B or Mistral Large better for coding?

Mistral Large scores higher on coding benchmarks: 34.3 versus 21.1 in the Noometry coding category.

Which has the bigger context window?

Mistral Large does, with 131K tokens against 60K.

How many benchmarks do Llama 3.2 1B and Mistral Large share?

20 benchmarks have published results for both models. Llama 3.2 1B has 22 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper