Model comparison

Claude 2.1 vs DeepSeek-V3.1

DeepSeek-V3.1 is the stronger model overall, scoring 42.8 to 25.2 on the Noometry Index.

Last verified . 4 shared benchmarks.

Claude 2.1 Anthropic

25.2

Rank #345 Reported

DeepSeek-V3.1 DeepSeek

42.8

Rank #108 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Claude 2.1 scores higher in 0 categories and DeepSeek-V3.1 in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in math, where DeepSeek-V3.1 leads 38.9 to 10.2.
  • The biggest single-benchmark swing is DTBench: 51% for Claude 2.1 and 82.7% for DeepSeek-V3.1.
  • DeepSeek-V3.1 has downloadable open weights; the other is API-only.

Side by side

Claude 2.1 and DeepSeek-V3.1 specifications
Claude 2.1DeepSeek-V3.1
ProviderAnthropicDeepSeek
Noometry Index25.242.8
Released2023-11-212025-08-21
WeightsProprietaryOpen
Context window—164K
Max output—8K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.95
Results tracked727

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3.1 leads

Claude 2.1: 26.2 (#327), DeepSeek-V3.1: 40.3 (#144)

Coding benchmarks
BenchmarkClaude 2.1DeepSeek-V3.1
WeirdML7.1%38.4%
LMArena Coding—1417

Reasoning DeepSeek-V3.1 leads

Claude 2.1: 21.4 (#221), DeepSeek-V3.1: 27.9 (#110)

Reasoning benchmarks
BenchmarkClaude 2.1DeepSeek-V3.1
DTBench51%82.7%
Epoch Capabilities Index119.27139.92
ForecastBench54.258
SimpleBench—40%
Kagi LLM Benchmark—53.2%
LMArena Hard Prompts—1417
LMCA—24.3%

Math DeepSeek-V3.1 leads

Claude 2.1: 10.2 (#315), DeepSeek-V3.1: 38.9 (#122)

Math benchmarks
BenchmarkClaude 2.1DeepSeek-V3.1
OTIS Mock AIME 2024-20251.9%—
LMArena Math—1420

Knowledge DeepSeek-V3.1 leads

Claude 2.1: 15.4 (#292), DeepSeek-V3.1: 43.7 (#90)

Knowledge benchmarks
BenchmarkClaude 2.1DeepSeek-V3.1
GPQA Diamond33%—
Vectara Hallucination Rate—5.5%
LMArena Expert—1405
MMLU73.5%—

Multilingual Not comparable

Claude 2.1: —, DeepSeek-V3.1: 51.6 (#106)

Multilingual benchmarks
BenchmarkClaude 2.1DeepSeek-V3.1
LMArena Non-English—1400
LMArena Chinese—1469
LMArena French—1447
LMArena German—1411
LMArena Japanese—1378
LMArena Korean—1337
LMArena Russian—1405
LMArena Spanish—1431

Instruction Following Not comparable

Claude 2.1: —, DeepSeek-V3.1: 73.9 (#110)

Instruction Following benchmarks
BenchmarkClaude 2.1DeepSeek-V3.1
LMArena Instruction Following—1400

Long Context Not comparable

Claude 2.1: —, DeepSeek-V3.1: 36.3 (#232)

Long Context benchmarks
BenchmarkClaude 2.1DeepSeek-V3.1
Fiction.LiveBench—52.8%
LMArena Longer Query—1422

Writing & Preference Not comparable

Claude 2.1: —, DeepSeek-V3.1: 60.3 (#98)

Writing & Preference benchmarks
BenchmarkClaude 2.1DeepSeek-V3.1
LMArena Text—1420
LMArena Creative Writing—1401
EQ-Bench Creative Writing—1436
LMArena Multi-Turn—1408

Frequently asked questions

Is Claude 2.1 better than DeepSeek-V3.1?

DeepSeek-V3.1 is the stronger model overall, scoring 42.8 to 25.2 on the Noometry Index.

Is Claude 2.1 or DeepSeek-V3.1 better for coding?

DeepSeek-V3.1 scores higher on coding benchmarks: 40.3 versus 26.2 in the Noometry coding category.

How many benchmarks do Claude 2.1 and DeepSeek-V3.1 share?

4 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and DeepSeek-V3.1 has 27.

Related comparisons

Go deeper