Model comparison

Falcon-180B vs Magistral Small

Falcon-180B is the stronger model overall, scoring 32.2 to 30.2 on the Noometry Index.

Last verified . 1 shared benchmarks.

Magistral Small Mistral AI

30.2

Rank #296 Confirmed

Summary

  • They share 1 benchmark with published results for both. Falcon-180B scores higher in 1 category and Magistral Small in 0 categories; one gap is clear of the uncertainty.
  • The widest gap is in reasoning, where Falcon-180B leads 19.1 to 6.8.

Side by side

Falcon-180B and Magistral Small specifications
Falcon-180BMagistral Small
ProviderTechnology Innovation InstituteMistral AI
Noometry Index32.230.2
Released2023-09-062025-06-10
WeightsOpenOpen
Context window—128K
Max output—40K
Input $ / M tokens—$0.50
Output $ / M tokens—$1.50
Results tracked1610

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Falcon-180B: —, Magistral Small: 38.4 (#176)

Coding benchmarks
BenchmarkFalcon-180BMagistral Small
SciCode—35.2%

Reasoning Falcon-180B leads

Falcon-180B: 19.1 (#269), Magistral Small: 6.8 (#350)

Reasoning benchmarks
BenchmarkFalcon-180BMagistral Small
Epoch Capabilities Index112.13133.19
ARC-AGI-2—0%
Kagi LLM Benchmark—6.3%
ARC-AGI-1—5%
CritPt—0.3%
Chess Puzzles—3%
LMArena Hard Prompts1007—
DTBench—61.3%
HellaSwag89%—
LAMBADA79.8%—
PIQA84.9%—
WinoGrande87.1%—

Math Not comparable

Falcon-180B: —, Magistral Small: 26.2 (#261)

Math benchmarks
BenchmarkFalcon-180BMagistral Small
OTIS Mock AIME 2024-2025—30%
GSM8K54.4%—

Knowledge Not comparable

Falcon-180B: —, Magistral Small: 30.9 (#223)

Knowledge benchmarks
BenchmarkFalcon-180BMagistral Small
GPQA Diamond—56.1%
ARC (AI2) Challenge67.8%—
BoolQ89%—
MMLU70.6%—
OpenBookQA64.2%—

Multilingual Not comparable

Falcon-180B: 25.2 (#286), Magistral Small: —

Multilingual benchmarks
BenchmarkFalcon-180BMagistral Small
LMArena Non-English1000—

Instruction Following Not comparable

Falcon-180B: 53.4 (#286), Magistral Small: —

Instruction Following benchmarks
BenchmarkFalcon-180BMagistral Small
LMArena Instruction Following1047—

Writing & Preference Not comparable

Falcon-180B: 29.1 (#295), Magistral Small: —

Writing & Preference benchmarks
BenchmarkFalcon-180BMagistral Small
LMArena Text1054—
LMArena Creative Writing1089—
LMArena Multi-Turn1013—

Frequently asked questions

Is Falcon-180B better than Magistral Small?

Falcon-180B is the stronger model overall, scoring 32.2 to 30.2 on the Noometry Index.

How many benchmarks do Falcon-180B and Magistral Small share?

1 benchmark has published results for both models. Falcon-180B has 16 scored results on Noometry and Magistral Small has 10.

Related comparisons

Go deeper