Model comparison

DeepSeek-V3.1 vs Falcon-180B

DeepSeek-V3.1 is the stronger model overall, scoring 42.8 to 32.2 on the Noometry Index.

Last verified . 7 shared benchmarks.

DeepSeek-V3.1 DeepSeek

42.8

Rank #108 Confirmed

Summary

  • They share 7 benchmarks with published results for both. DeepSeek-V3.1 scores higher in 4 categories and Falcon-180B in 0 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where DeepSeek-V3.1 leads 60.3 to 29.1.

Side by side

DeepSeek-V3.1 and Falcon-180B specifications
DeepSeek-V3.1Falcon-180B
ProviderDeepSeekTechnology Innovation Institute
Noometry Index42.832.2
Released2025-08-212023-09-06
WeightsOpenOpen
Context window164K—
Max output8K—
Input $ / M tokens$0.25—
Output $ / M tokens$0.95—
Results tracked2716

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

DeepSeek-V3.1: 40.3 (#144), Falcon-180B: —

Coding benchmarks
BenchmarkDeepSeek-V3.1Falcon-180B
WeirdML38.4%—
LMArena Coding1417—

Reasoning DeepSeek-V3.1 leads

DeepSeek-V3.1: 27.9 (#110), Falcon-180B: 19.1 (#269)

Reasoning benchmarks
BenchmarkDeepSeek-V3.1Falcon-180B
LMArena Hard Prompts14171007
Epoch Capabilities Index139.92112.13
SimpleBench40%—
Kagi LLM Benchmark53.2%—
DTBench82.7%—
LMCA24.3%—
ForecastBench58—
HellaSwag—89%
LAMBADA—79.8%
PIQA—84.9%
WinoGrande—87.1%

Math Not comparable

DeepSeek-V3.1: 38.9 (#122), Falcon-180B: —

Math benchmarks
BenchmarkDeepSeek-V3.1Falcon-180B
LMArena Math1420—
GSM8K—54.4%

Knowledge Not comparable

DeepSeek-V3.1: 43.7 (#90), Falcon-180B: —

Knowledge benchmarks
BenchmarkDeepSeek-V3.1Falcon-180B
Vectara Hallucination Rate5.5%—
LMArena Expert1405—
ARC (AI2) Challenge—67.8%
BoolQ—89%
MMLU—70.6%
OpenBookQA—64.2%

Multilingual DeepSeek-V3.1 leads

DeepSeek-V3.1: 51.6 (#106), Falcon-180B: 25.2 (#286)

Multilingual benchmarks
BenchmarkDeepSeek-V3.1Falcon-180B
LMArena Non-English14001000
LMArena Chinese1469—
LMArena French1447—
LMArena German1411—
LMArena Japanese1378—
LMArena Korean1337—
LMArena Russian1405—
LMArena Spanish1431—

Instruction Following DeepSeek-V3.1 leads

DeepSeek-V3.1: 73.9 (#110), Falcon-180B: 53.4 (#286)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.1Falcon-180B
LMArena Instruction Following14001047

Long Context Not comparable

DeepSeek-V3.1: 36.3 (#232), Falcon-180B: —

Long Context benchmarks
BenchmarkDeepSeek-V3.1Falcon-180B
Fiction.LiveBench52.8%—
LMArena Longer Query1422—

Writing & Preference DeepSeek-V3.1 leads

DeepSeek-V3.1: 60.3 (#98), Falcon-180B: 29.1 (#295)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.1Falcon-180B
LMArena Text14201054
LMArena Creative Writing14011089
LMArena Multi-Turn14081013
EQ-Bench Creative Writing1436—

Frequently asked questions

Is DeepSeek-V3.1 better than Falcon-180B?

DeepSeek-V3.1 is the stronger model overall, scoring 42.8 to 32.2 on the Noometry Index.

How many benchmarks do DeepSeek-V3.1 and Falcon-180B share?

7 benchmarks have published results for both models. DeepSeek-V3.1 has 27 scored results on Noometry and Falcon-180B has 16.

Related comparisons

Go deeper