Model comparison

Falcon-180B vs Mistral Small 3.1

Falcon-180B and Mistral Small 3.1 score almost the same on the Noometry Index (32.2 vs 31.7), so choose on price, context window or the category you care about most.

Last verified . 7 shared benchmarks.

Mistral Small 3.1 Mistral AI

31.7

Rank #269 Confirmed

Summary

  • They share 7 benchmarks with published results for both. Falcon-180B scores higher in 0 categories and Mistral Small 3.1 in 4 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Mistral Small 3.1 leads 41.2 to 25.2.

Side by side

Falcon-180B and Mistral Small 3.1 specifications
Falcon-180BMistral Small 3.1
ProviderTechnology Innovation InstituteMistral AI
Noometry Index32.231.7
Released2023-09-062025-03-17
WeightsOpenOpen
Context window—128K
Max output—102K
Input $ / M tokens—$0.35
Output $ / M tokens—$0.56
Results tracked1628

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Falcon-180B: —, Mistral Small 3.1: 38.3 (#179)

Coding benchmarks
BenchmarkFalcon-180BMistral Small 3.1
LMArena Coding—1309

Reasoning Too close to call

Falcon-180B: 19.1 (#269), Mistral Small 3.1: 19.7 (#254)

Reasoning benchmarks
BenchmarkFalcon-180BMistral Small 3.1
LMArena Hard Prompts10071278
Epoch Capabilities Index112.13127.48
Chess Puzzles—1%
HellaSwag89%—
LAMBADA79.8%—
PIQA84.9%—
WinoGrande87.1%—

Math Not comparable

Falcon-180B: —, Mistral Small 3.1: 14.7 (#301)

Math benchmarks
BenchmarkFalcon-180BMistral Small 3.1
OTIS Mock AIME 2024-2025—3.9%
Omni-MATH—24.8%
LMArena Math—1262
GSM8K54.4%—

Knowledge Not comparable

Falcon-180B: —, Mistral Small 3.1: 22.6 (#271)

Knowledge benchmarks
BenchmarkFalcon-180BMistral Small 3.1
GPQA Diamond—41.9%
MMLU-Pro—61%
GPQA (HELM)—39.2%
LMArena Expert—1257
ARC (AI2) Challenge67.8%—
BoolQ89%—
MMLU70.6%—
OpenBookQA64.2%—

Multimodal Not comparable

Falcon-180B: —, Mistral Small 3.1: 33.2 (#99)

Multimodal benchmarks
BenchmarkFalcon-180BMistral Small 3.1
LMArena Vision—1136

Multilingual Mistral Small 3.1 leads

Falcon-180B: 25.2 (#286), Mistral Small 3.1: 41.2 (#209)

Multilingual benchmarks
BenchmarkFalcon-180BMistral Small 3.1
LMArena Non-English10001255
LMArena Chinese—1253
LMArena French—1273
LMArena German—1266
LMArena Japanese—1208
LMArena Korean—1206
LMArena Russian—1263
LMArena Spanish—1283

Instruction Following Mistral Small 3.1 leads

Falcon-180B: 53.4 (#286), Mistral Small 3.1: 63.6 (#230)

Instruction Following benchmarks
BenchmarkFalcon-180BMistral Small 3.1
LMArena Instruction Following10471264
IFEval—75%

Long Context Not comparable

Falcon-180B: —, Mistral Small 3.1: 39.5 (#178)

Long Context benchmarks
BenchmarkFalcon-180BMistral Small 3.1
LMArena Longer Query—1299

Writing & Preference Mistral Small 3.1 leads

Falcon-180B: 29.1 (#295), Mistral Small 3.1: 37.0 (#259)

Writing & Preference benchmarks
BenchmarkFalcon-180BMistral Small 3.1
LMArena Text10541277
LMArena Creative Writing10891253
LMArena Multi-Turn10131270
EQ-Bench Creative Writing—761
WildBench—78.8%

Frequently asked questions

Is Falcon-180B better than Mistral Small 3.1?

Falcon-180B and Mistral Small 3.1 score almost the same on the Noometry Index (32.2 vs 31.7), so choose on price, context window or the category you care about most.

How many benchmarks do Falcon-180B and Mistral Small 3.1 share?

7 benchmarks have published results for both models. Falcon-180B has 16 scored results on Noometry and Mistral Small 3.1 has 28.

Related comparisons

Go deeper