DeepSeek, proprietary

DeepSeek-R1

DeepSeek-R1 by DeepSeek ranks 115th of 354 ranked models on the Noometry Index as of October 2026, with a score of 42.3. Its strongest category is long context, where it ranks 36th. API pricing starts at $0.50 per million input tokens and $2.15 per million output tokens, with a 164K-token context window.

Last verified

Specifications

Noometry rank
#115 of 354
Index score
42.3
Evidence
Confirmed 52 results
Provider
DeepSeek
Released
January 20, 2025
Weights
Proprietary
Reasoning
Yes
Context window
164K
Max output
64K
Input price
$0.50 / M
Output price
$2.15 / M
Blended price
$0.91 / M
Output speed
10 tokens/s Kagi
Value
#101 of 219
Knowledge cutoff
July 2024
Input
text

Category scores

Each category score combines every public result we have in that category.

DeepSeek-R1 category scores
  1. Coding 46.3
  2. Agentic & Tool Use 30.7
  3. Reasoning 18.6
  4. Math 43.8
  5. Knowledge 44.5
  6. Multilingual 52.4
  7. Instruction Following 72.0
  8. Long Context 45.4
  9. Writing & Preference 61.4
DeepSeek-R1 category ranks
CategoryScoreRankResults
Coding46.3#685
Agentic & Tool Use30.7#752
Reasoning18.6#2788
Math43.8#795
Knowledge44.5#876
Multilingual52.4#851
Instruction Following72.0#1433
Long Context45.4#362
Writing & Preference61.4#887

Strengths and weaknesses

Categories where DeepSeek-R1 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

DeepSeek-R1: strongest categories
CategoryScorevs medianRank
Long Context45.4+4.5#36 of 296, top 13%
Coding46.3+7.6#68 of 340, top 20%
Math43.8+7.2#79 of 327, top 25%

Weakest categories

DeepSeek-R1: weakest categories
CategoryScorevs medianRank
Reasoning18.6−5.0#278 of 350, top 80%
Agentic & Tool Use30.7+0.3#75 of 154, top 49%
Instruction Following72.0+0.7#143 of 305, top 47%

Closest competitors

The models ranked just above and below DeepSeek-R1. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to DeepSeek-R1
ModelRankScoreBlended $/MSpeed
GPT-5.2 Codex#11142.6$4.81—Compare
Qwen3.5-Flash#11242.5$0.18—Compare
Nemotron 3 Ultra#11342.5$0.93—Compare
Hunyuan T1 20250711#11442.5——Compare
Step 3.5 Flash#11642.3$0.15—Compare
Qwen3.6 27B#11742.2$1.35—Compare
Amazon Nova Experimental Chat 10 20#11842.1——Compare
Qwen3.5 122B-A10B#11942.1$1.10—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

DeepSeek-R1 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
Aider Polyglot56.9%Epoch AI
Aider Polyglot71.4%#9 of 44, top 21%Epoch AI
SciCode35.7%#97 of 121, top 81%Epoch AI
WeirdML36.5%Epoch AI
WeirdML41.6%#72 of 119, top 61%Epoch AI
LiveBench Coding66.7%#10 of 39, top 26%Epoch AI
LMArena Coding1427#112 of 294, top 39%LMArena2026-10-08
LMArena Coding1371LMArena2026-10-08
ALE-Bench804.12#56 of 105, top 54%Epoch AI
AlgoTune1.7#7 of 18, top 39%Epoch AI

Agentic & Tool Use

DeepSeek-R1 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
DeepResearch Bench35.1%#23 of 24, top 96%Epoch AI
BALROG34.9%#12 of 35, top 35%Epoch AI
METR Time Horizons51.9%Epoch AI
METR Time Horizons53.8%#22 of 32, top 69%Epoch AI

Reasoning

DeepSeek-R1 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-21.3%#66 of 83, top 80%Epoch AI
ARC-AGI-21.1%Epoch AI
SimpleBench40.8%#53 of 77, top 69%Epoch AI
SimpleBench30.9%Epoch AI
Kagi LLM Benchmark69.4%#25 of 99, top 26%Kagi LLM Benchmark
ARC-AGI-115.8%Epoch AI
ARC-AGI-121.2%#68 of 83, top 82%Epoch AI
CritPt1.1%#82 of 134, top 62%Epoch AI
LiveBench Reasoning83.2%#7 of 39, top 18%Epoch AI
LMArena Hard Prompts1416#111 of 297, top 38%LMArena2026-10-08
LMArena Hard Prompts1361LMArena2026-10-08
LiveBench Data Analysis69.8%#5 of 39, top 13%Epoch AI
Epoch Capabilities Index141.29#104 of 213, top 49%Epoch AI2025-05-28
Epoch Capabilities Index138.97Epoch AI2025-01-20
ForecastBench60#33 of 72, top 46%Epoch AI
LiveBench71.6%#7 of 39, top 18%Epoch AI

Math

DeepSeek-R1 Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-202566.4%#98 of 173, top 57%Epoch AI2025-05-29
OTIS Mock AIME 2024-202553.3%Epoch AI2025-02-26
Omni-MATH42.4%#23 of 57, top 41%HELM Capabilities
LiveBench Math80.7%#3 of 39, top 8%Epoch AI
LMArena Math1393LMArena2026-10-08
LMArena Math1400#125 of 285, top 44%LMArena2026-10-08
MATH Level 596.6%#7 of 79, top 9%Epoch AI2025-05-29
MATH Level 593.1%Epoch AI2025-01-31

Knowledge

DeepSeek-R1 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond76.3%#88 of 186, top 48%Epoch AI2025-05-29
GPQA Diamond71.7%Epoch AI2025-05-26
MMLU-Pro79.3%#18 of 58, top 32%HELM Capabilities
Confabulations (lower is better)12.7%#9 of 51, top 18%Lech Mazur benchmarks
Confabulations (lower is better)14.6%Lech Mazur benchmarks
Vectara Hallucination Rate (lower is better)11.3%#68 of 96, top 71%Vectara Hallucination Leaderboard
GPQA (HELM)66.6%#15 of 57, top 27%HELM Capabilities
LMArena Expert1394#131 of 273, top 48%LMArena2026-10-08
LMArena Expert1338LMArena2026-10-08

Multilingual

DeepSeek-R1 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1412#85 of 297, top 29%LMArena2026-10-08
LMArena Non-English1357LMArena2026-10-08
LMArena Chinese1400LMArena2026-10-08
LMArena Chinese1442#110 of 285, top 39%LMArena2026-10-08
LMArena French1417#105 of 223, top 48%LMArena2026-10-08
LMArena French1366LMArena2026-10-08
LMArena German1404#92 of 231, top 40%LMArena2026-10-08
LMArena German1384LMArena2026-10-08
LMArena Japanese1324LMArena2026-10-08
LMArena Japanese1391#68 of 211, top 33%LMArena2026-10-08
LMArena Korean1329LMArena2026-10-08
LMArena Korean1360#94 of 213, top 45%LMArena2026-10-08
LMArena Russian1354LMArena2026-10-08
LMArena Russian1423#76 of 283, top 27%LMArena2026-10-08
LMArena Spanish1411#100 of 226, top 45%LMArena2026-10-08
LMArena Spanish1379LMArena2026-10-08

Instruction Following

DeepSeek-R1 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LiveBench Instruction Following80.5%#11 of 39, top 29%Epoch AI
IFEval78.4%#45 of 57, top 79%HELM Capabilities
LMArena Instruction Following1357LMArena2026-10-08
LMArena Instruction Following1382#120 of 298, top 41%LMArena2026-10-08

Long Context

DeepSeek-R1 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
Fiction.LiveBench69.4%Epoch AI
Fiction.LiveBench75%#14 of 47, top 30%Epoch AI
LMArena Longer Query1391#125 of 291, top 43%LMArena2026-10-08
LMArena Longer Query1355LMArena2026-10-08

Writing & Preference

DeepSeek-R1 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1428#81 of 297, top 28%LMArena2026-10-08
LMArena Text1373LMArena2026-10-08
LMArena Creative Writing1354LMArena2026-10-08
LMArena Creative Writing1405#68 of 295, top 24%LMArena2026-10-08
Short-Story Creative Writing81.9%Epoch AI
Short-Story Creative Writing83%#9 of 39, top 24%Epoch AI
EQ-Bench Creative Writing1421EQ-Bench
EQ-Bench Creative Writing1500#55 of 115, top 48%EQ-Bench
WildBench82.8%#20 of 57, top 36%HELM Capabilities
LMArena Multi-Turn1391LMArena2026-10-08
LMArena Multi-Turn1405#118 of 295, top 40%LMArena2026-10-08
LiveBench Language48.5%#13 of 39, top 34%Epoch AI

API pricing by provider

DeepSeek-R1 API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$1.35$5.40—2026-10-10
bedrock$1.35$5.40—2026-10-10
deepinfra$0.50$2.15$0.352026-10-10
openrouter$0.70$2.50—2026-10-10
together$3$7—2026-10-10

Compare DeepSeek-R1

Other DeepSeek models

Frequently asked questions

How good is DeepSeek-R1?

DeepSeek-R1 by DeepSeek ranks 115th of 354 ranked models on the Noometry Index as of October 2026, with a score of 42.3. Its strongest category is long context, where it ranks 36th. API pricing starts at $0.50 per million input tokens and $2.15 per million output tokens, with a 164K-token context window.

How much does DeepSeek-R1 cost?

DeepSeek-R1 costs $0.50 per million input tokens and $2.15 per million output tokens on deepinfra, with cached input at $0.35.

What is DeepSeek-R1's context window?

DeepSeek-R1 accepts up to 164K tokens of input and can write up to 64K tokens in one response.

Is DeepSeek-R1 open source?

No. DeepSeek-R1 is proprietary and available only through DeepSeek's API and partner platforms.

How fast is DeepSeek-R1?

DeepSeek-R1 generated about 10 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are DeepSeek-R1's strengths and weaknesses?

Relative to other ranked models, DeepSeek-R1 places best in long context, coding, math and lowest in reasoning, agentic & tool use, instruction following.

What is DeepSeek-R1 best at?

Its best category is long context, where it ranks 36th on Noometry.