Model comparison

Llama 3.2 1B vs Mistral Medium 3.1

Mistral Medium 3.1 is the stronger model overall, scoring 31.9 to 20.1 on the Noometry Index. Llama 3.2 1B costs 11× less per token, which makes it the better buy when Mistral Medium 3.1's lead doesn't matter for your workload.

Last verified . 1 shared benchmarks.

Llama 3.2 1B Meta

20.1

Rank #354 Confirmed

Mistral Medium 3.1 Mistral AI

31.9

Rank #266 Reported

Summary

  • They share 1 benchmark with published results for both. Llama 3.2 1B scores higher in 1 category and Mistral Medium 3.1 in 1 category; 2 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Medium 3.1 leads 55.5 to 21.3.
  • Llama 3.2 1B is cheaper at $0.027 / $0.20 per million input/output tokens, against $0.40 / $2 for Mistral Medium 3.1.
  • Mistral Medium 3.1 accepts more context: 131K tokens versus 60K.
  • Llama 3.2 1B has downloadable open weights; the other is API-only.

Side by side

Llama 3.2 1B and Mistral Medium 3.1 specifications
Llama 3.2 1BMistral Medium 3.1
ProviderMetaMistral AI
Noometry Index20.131.9
Released2024-09-24—
WeightsOpenProprietary
Context window60K131K
Max output54K105K
Input $ / M tokens$0.027$0.40
Output $ / M tokens$0.20$2
Results tracked223

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 3.2 1B: 21.1 (#338), Mistral Medium 3.1: —

Coding benchmarks
BenchmarkLlama 3.2 1BMistral Medium 3.1
BigCodeBench Instruct8.2%—
LMArena Coding1070—
BigCodeBench Complete11.3%—

Agentic & Tool Use Not comparable

Llama 3.2 1B: 14.6 (#150), Mistral Medium 3.1: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 1BMistral Medium 3.1
Berkeley Function Calling Leaderboard10.8%—
BALROG6.6%—

Reasoning Llama 3.2 1B leads

Llama 3.2 1B: 16.2 (#308), Mistral Medium 3.1: 10.6 (#341)

Reasoning benchmarks
BenchmarkLlama 3.2 1BMistral Medium 3.1
NYT Connections (extended)—6.5%
Chess Puzzles0%—
Thematic Generalization—20.3%
LMArena Hard Prompts1044—
Epoch Capabilities Index101.99—

Math Not comparable

Llama 3.2 1B: 10.4 (#313), Mistral Medium 3.1: —

Math benchmarks
BenchmarkLlama 3.2 1BMistral Medium 3.1
OTIS Mock AIME 2024-20250.6%—
LMArena Math1086—

Knowledge Not comparable

Llama 3.2 1B: 7.2 (#312), Mistral Medium 3.1: —

Knowledge benchmarks
BenchmarkLlama 3.2 1BMistral Medium 3.1
GPQA Diamond23.9%—
LMArena Expert1007—

Multilingual Not comparable

Llama 3.2 1B: 23.8 (#292), Mistral Medium 3.1: —

Multilingual benchmarks
BenchmarkLlama 3.2 1BMistral Medium 3.1
LMArena Non-English973—
LMArena Chinese959—
LMArena German1014—
LMArena Russian941—

Instruction Following Not comparable

Llama 3.2 1B: 52.4 (#290), Mistral Medium 3.1: —

Instruction Following benchmarks
BenchmarkLlama 3.2 1BMistral Medium 3.1
LMArena Instruction Following1031—

Long Context Not comparable

Llama 3.2 1B: 31.9 (#274), Mistral Medium 3.1: —

Long Context benchmarks
BenchmarkLlama 3.2 1BMistral Medium 3.1
LMArena Longer Query1050—

Writing & Preference Mistral Medium 3.1 leads

Llama 3.2 1B: 21.3 (#310), Mistral Medium 3.1: 55.5 (#145)

Writing & Preference benchmarks
BenchmarkLlama 3.2 1BMistral Medium 3.1
EQ-Bench Creative Writing2001476
LMArena Text1055—
LMArena Creative Writing1033—
LMArena Multi-Turn1030—

Frequently asked questions

Is Llama 3.2 1B better than Mistral Medium 3.1?

Mistral Medium 3.1 is the stronger model overall, scoring 31.9 to 20.1 on the Noometry Index. Llama 3.2 1B costs 11× less per token, which makes it the better buy when Mistral Medium 3.1's lead doesn't matter for your workload.

Which is cheaper, Llama 3.2 1B or Mistral Medium 3.1?

Llama 3.2 1B is cheaper. It lists at $0.027 per million input tokens and $0.20 per million output tokens; Mistral Medium 3.1 lists at $0.40 and $2.

Which has the bigger context window?

Mistral Medium 3.1 does, with 131K tokens against 60K.

How many benchmarks do Llama 3.2 1B and Mistral Medium 3.1 share?

1 benchmark has published results for both models. Llama 3.2 1B has 22 scored results on Noometry and Mistral Medium 3.1 has 3.

Related comparisons

Go deeper