Model comparison

DeepSeek-V3.2-Speciale vs Falcon-180B

DeepSeek-V3.2-Speciale is the stronger model overall, scoring 39.7 to 32.2 on the Noometry Index.

Last verified . 0 shared benchmarks.

Summary

  • The widest gap is in writing & preference, where DeepSeek-V3.2-Speciale leads 46.0 to 29.1.

Side by side

DeepSeek-V3.2-Speciale and Falcon-180B specifications
DeepSeek-V3.2-SpecialeFalcon-180B
ProviderDeepSeekTechnology Innovation Institute
Noometry Index39.732.2
Released2025-12-012023-09-06
WeightsOpenOpen
Context window128K—
Max output128K—
Input $ / M tokens$0.58—
Output $ / M tokens$1.68—
Results tracked316

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

DeepSeek-V3.2-Speciale: 40.4 (#140), Falcon-180B: —

Coding benchmarks
BenchmarkDeepSeek-V3.2-SpecialeFalcon-180B
WeirdML46.7%—

Reasoning DeepSeek-V3.2-Speciale leads

DeepSeek-V3.2-Speciale: 32.9 (#73), Falcon-180B: 19.1 (#269)

Reasoning benchmarks
BenchmarkDeepSeek-V3.2-SpecialeFalcon-180B
SimpleBench52.6%—
LMArena Hard Prompts—1007
Epoch Capabilities Index—112.13
HellaSwag—89%
LAMBADA—79.8%
PIQA—84.9%
WinoGrande—87.1%

Math Not comparable

DeepSeek-V3.2-Speciale: —, Falcon-180B: —

Math benchmarks
BenchmarkDeepSeek-V3.2-SpecialeFalcon-180B
GSM8K—54.4%

Knowledge Not comparable

DeepSeek-V3.2-Speciale: —, Falcon-180B: —

Knowledge benchmarks
BenchmarkDeepSeek-V3.2-SpecialeFalcon-180B
ARC (AI2) Challenge—67.8%
BoolQ—89%
MMLU—70.6%
OpenBookQA—64.2%

Multilingual Not comparable

DeepSeek-V3.2-Speciale: —, Falcon-180B: 25.2 (#286)

Multilingual benchmarks
BenchmarkDeepSeek-V3.2-SpecialeFalcon-180B
LMArena Non-English—1000

Instruction Following Not comparable

DeepSeek-V3.2-Speciale: —, Falcon-180B: 53.4 (#286)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.2-SpecialeFalcon-180B
LMArena Instruction Following—1047

Writing & Preference DeepSeek-V3.2-Speciale leads

DeepSeek-V3.2-Speciale: 46.0 (#222), Falcon-180B: 29.1 (#295)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.2-SpecialeFalcon-180B
LMArena Text—1054
LMArena Creative Writing—1089
EQ-Bench Creative Writing1276—
LMArena Multi-Turn—1013

Frequently asked questions

Is DeepSeek-V3.2-Speciale better than Falcon-180B?

DeepSeek-V3.2-Speciale is the stronger model overall, scoring 39.7 to 32.2 on the Noometry Index.

How many benchmarks do DeepSeek-V3.2-Speciale and Falcon-180B share?

0 benchmarks have published results for both models. DeepSeek-V3.2-Speciale has 3 scored results on Noometry and Falcon-180B has 16.

Related comparisons

Go deeper