Model comparison

Ministral 3B vs Mistral Large

Mistral Large is the stronger model overall, scoring 31.9 to 26.2 on the Noometry Index. Ministral 3B costs 30× less per token, which makes it the better buy when Mistral Large's lead doesn't matter for your workload.

Last verified . 6 shared benchmarks.

Ministral 3B Mistral AI

26.2

Rank #338 Confirmed

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Summary

  • They share 6 benchmarks with published results for both. Ministral 3B scores higher in 2 categories and Mistral Large in 1 category; 3 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Mistral Large leads 30.1 to 10.4.
  • The biggest single-benchmark swing is MATH Level 5: 14.4% for Ministral 3B and 50.3% for Mistral Large.
  • Ministral 3B is cheaper at $0.10 / $0.10 per million input/output tokens, against $2 / $6 for Mistral Large.

Side by side

Ministral 3B and Mistral Large specifications
Ministral 3BMistral Large
ProviderMistral AIMistral AI
Noometry Index26.231.9
Released2024-10-012024-02-26
WeightsOpenOpen
Context window131K131K
Max output262K16K
Input $ / M tokens$0.10$2
Output $ / M tokens$0.10$6
Results tracked651

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Ministral 3B: —, Mistral Large: 34.3 (#240)

Coding benchmarks
BenchmarkMinistral 3BMistral Large
SciCode—36.2%
BigCodeBench Instruct—30%
LiveBench Coding—47.1%
LMArena Coding—1277
BigCodeBench Complete—38.3%
ALE-Bench—264.7
HumanEval+—62.2%
MBPP+—59.5%

Agentic & Tool Use Not comparable

Ministral 3B: —, Mistral Large: 28.6 (#89)

Agentic & Tool Use benchmarks
BenchmarkMinistral 3BMistral Large
Berkeley Function Calling Leaderboard—38.4%

Reasoning Ministral 3B leads

Ministral 3B: 18.4 (#282), Mistral Large: 15.8 (#310)

Reasoning benchmarks
BenchmarkMinistral 3BMistral Large
DTBench51.7%65.1%
LMCA5.5%16.7%
Epoch Capabilities Index118.1128.52
SimpleBench—22.5%
CritPt—0%
LiveBench Reasoning—43.5%
LMArena Hard Prompts—1257
LiveBench Data Analysis—50.1%
ForecastBench—57.1
LiveBench—48.4%

Math Ministral 3B leads

Ministral 3B: 26.6 (#258), Mistral Large: 18.2 (#291)

Math benchmarks
BenchmarkMinistral 3BMistral Large
MATH Level 514.4%50.3%
OTIS Mock AIME 2024-2025—8.5%
Omni-MATH—28.1%
LiveBench Math—42.5%
LMArena Math—1262
FrontierMath (Feb 2025 set)—0.3%

Knowledge Mistral Large leads

Ministral 3B: 10.4 (#302), Mistral Large: 30.1 (#230)

Knowledge benchmarks
BenchmarkMinistral 3BMistral Large
GPQA Diamond25.3%51.3%
Vectara Hallucination Rate7.3%4.5%
MMLU-Pro—59.9%
Confabulations—21.4%
GPQA (HELM)—43.5%
LMArena Expert—1232
MMLU—80%

Multilingual Not comparable

Ministral 3B: —, Mistral Large: 40.0 (#219)

Multilingual benchmarks
BenchmarkMinistral 3BMistral Large
LMArena Non-English—1237
LMArena Chinese—1240
LMArena French—1325
LMArena German—1254
LMArena Japanese—1188
LMArena Korean—1202
LMArena Russian—1257
LMArena Spanish—1268

Instruction Following Not comparable

Ministral 3B: —, Mistral Large: 67.9 (#191)

Instruction Following benchmarks
BenchmarkMinistral 3BMistral Large
LiveBench Instruction Following—67.9%
IFEval—87.7%
LMArena Instruction Following—1249

Long Context Not comparable

Ministral 3B: —, Mistral Large: 38.3 (#199)

Long Context benchmarks
BenchmarkMinistral 3BMistral Large
LMArena Longer Query—1261

Writing & Preference Not comparable

Ministral 3B: —, Mistral Large: 40.7 (#242)

Writing & Preference benchmarks
BenchmarkMinistral 3BMistral Large
LMArena Text—1266
LMArena Creative Writing—1243
Short-Story Creative Writing—69%
EQ-Bench Creative Writing—985
WildBench—80.1%
LMArena Multi-Turn—1260
LiveBench Language—39.4%

Frequently asked questions

Is Ministral 3B better than Mistral Large?

Mistral Large is the stronger model overall, scoring 31.9 to 26.2 on the Noometry Index. Ministral 3B costs 30× less per token, which makes it the better buy when Mistral Large's lead doesn't matter for your workload.

Which is cheaper, Ministral 3B or Mistral Large?

Ministral 3B is cheaper. It lists at $0.10 per million input tokens and $0.10 per million output tokens; Mistral Large lists at $2 and $6.

Which has the bigger context window?

Both accept 131K tokens.

How many benchmarks do Ministral 3B and Mistral Large share?

6 benchmarks have published results for both models. Ministral 3B has 6 scored results on Noometry and Mistral Large has 51.

Related comparisons

Go deeper