DeepSeek, open weights

DeepSeek V4 Pro

DeepSeek V4 Pro by DeepSeek ranks 31st of 354 ranked models on the Noometry Index as of October 2026, with a score of 54.3. Its strongest category is reasoning, where it ranks 24th. API pricing starts at $0.66 per million input tokens and $1.98 per million output tokens, with a 1M-token context window.

Last verified

Specifications

Noometry rank
#31 of 354
Index score
54.3
Evidence
Confirmed 48 results
Provider
DeepSeek
Released
April 24, 2026
Weights
Open weights
Reasoning
Yes
Context window
1M
Max output
393K
Input price
$0.66 / M
Output price
$1.98 / M
Blended price
$0.99 / M
Output speed
16 tokens/s Kagi
Value
#92 of 219
Knowledge cutoff
May 2025
Input
text

Category scores

Each category score combines every public result we have in that category.

DeepSeek V4 Pro category scores
  1. Coding 52.4
  2. Agentic & Tool Use 32.8
  3. Reasoning 56.5
  4. Math 64.8
  5. Knowledge 59.5
  6. Multilingual 54.4
  7. Instruction Following 76.1
  8. Long Context 45.0
  9. Writing & Preference 65.5
DeepSeek V4 Pro category ranks
CategoryScoreRankResults
Coding52.4#346
Agentic & Tool Use32.8#581
Reasoning56.5#2411
Math64.8#306
Knowledge59.5#314
Multilingual54.4#451
Instruction Following76.1#471
Long Context45.0#512
Writing & Preference65.5#465

Strengths and weaknesses

Categories where DeepSeek V4 Pro places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

DeepSeek V4 Pro: strongest categories
CategoryScorevs medianRank
Reasoning56.5+32.9#24 of 350, top 7%
Math64.8+28.3#30 of 327, top 10%
Knowledge59.5+22.2#31 of 314, top 10%

Weakest categories

DeepSeek V4 Pro: weakest categories
CategoryScorevs medianRank
Agentic & Tool Use32.8+2.5#58 of 154, top 38%
Long Context45.0+4.0#51 of 296, top 18%
Instruction Following76.1+4.9#47 of 305, top 16%

Closest competitors

The models ranked just above and below DeepSeek V4 Pro. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to DeepSeek V4 Pro
ModelRankScoreBlended $/MSpeed
Muse Spark 1.3#2754.8$2—Compare
Gemini 3 Pro#2854.8—1Compare
Claude Sonnet 5#2954.6$4—Compare
GPT-5.6 Luna#3054.6$0.4512Compare
Gemini 3.5 Flash#3254.2$3.38—Compare
Gemini 3.6 Flash#3354.1$1.50—Compare
GPT-5.2#3454.1$4.8115Compare
DeepSeek V4 Flash#3553.6$0.266Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

DeepSeek V4 Pro Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified77.6%#6 of 32, top 19%maxEpoch AI2026-06-18
FrontierCode28.6%#27 of 37, top 73%highEpoch AI
FrontierCode17.6%noneEpoch AI
LMArena WebDev1464LMArena2026-10-08
LMArena WebDev1582#27 of 113, top 24%LMArena2026-10-08
LMArena WebDev1447LMArena2026-10-08
SciCode46.4%highEpoch AI
SciCode50%maxEpoch AI
SciCode51%#40 of 121, top 34%maxEpoch AI
SciCode39.9%noneEpoch AI
WeirdML46.5%highEpoch AI
WeirdML66.2%#22 of 119, top 19%maxEpoch AI
WeirdML48.9%maxEpoch AI
LMArena Coding1469LMArena2026-10-08
LMArena Coding1470#57 of 294, top 20%LMArena2026-10-08
LMArena Coding1453LMArena2026-10-08
ALE-Bench1,006highEpoch AI
ALE-Bench1,403#19 of 105, top 19%maxEpoch AI
ALE-Bench521.67noneEpoch AI

Agentic & Tool Use

DeepSeek V4 Pro Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
APEX-Agents47.3%#29 of 49, top 60%Epoch AI
Vending-Bench 23,285#41 of 60, top 69%Epoch AI

Reasoning

DeepSeek V4 Pro Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-259.7%highEpoch AI
ARC-AGI-256.3%lowEpoch AI
ARC-AGI-261.3%#26 of 83, top 32%maxEpoch AI
ARC-AGI-20.8%noneEpoch AI
Kagi LLM Benchmark53.5%#54 of 99, top 55%Kagi LLM Benchmark
NYT Connections (extended)67.3%Lech Mazur benchmarks
NYT Connections (extended)91.3%#14 of 91, top 16%high reasoningLech Mazur benchmarks
ARC-AGI-187.2%highEpoch AI
ARC-AGI-190.5%#25 of 83, top 31%lowEpoch AI
ARC-AGI-190%maxEpoch AI
ARC-AGI-113%noneEpoch AI
CritPt10%highEpoch AI
CritPt12.9%maxEpoch AI
CritPt18%#29 of 134, top 22%maxEpoch AI
CritPt0%noneEpoch AI
Chess Puzzles13%highEpoch AI2026-08-06
Chess Puzzles47%#13 of 129, top 11%maxEpoch AI2026-08-18
Chess Puzzles20%maxEpoch AI2026-06-16
Chess Puzzles7%noneEpoch AI2026-08-06
LMArena Hard Prompts1452LMArena2026-10-08
LMArena Hard Prompts1458LMArena2026-10-08
LMArena Hard Prompts1461#46 of 297, top 16%LMArena2026-10-08
Mystery Game Puzzles43%#13 of 74, top 18%maxEpoch AI2026-08-19
Mystery Game Puzzles17%noneEpoch AI2026-08-27
DTBench93.9%#24 of 151, top 16%Epoch AI
DTBench90.7%maxEpoch AI
LMCA45.5%#32 of 125, top 26%Epoch AI
LMCA41.2%maxEpoch AI
Surface Evolver Bench40%#20 of 25, top 80%highEpoch AI
Epoch Capabilities Index149.13Epoch AI2026-04-24
Epoch Capabilities Index155.31#29 of 213, top 14%Epoch AI2026-08-13
ForecastBench56.1#66 of 72, top 92%Epoch AI

Math

DeepSeek V4 Pro Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)45.3%maxEpoch AI2026-06-17
FrontierMath (Tiers 1-3)64.6%#32 of 81, top 40%maxEpoch AI2026-08-19
FrontierMath Tier 42.4%maxEpoch AI2026-06-17
FrontierMath Tier 426.8%#34 of 63, top 54%maxEpoch AI2026-08-19
MathArena Final-Answer Competitions76.6%#7 of 29, top 25%maxMathArena
OTIS Mock AIME 2024-202595.6%highEpoch AI2026-08-06
OTIS Mock AIME 2024-202596.7%maxEpoch AI2026-06-17
OTIS Mock AIME 2024-202598.6%#19 of 173, top 11%maxEpoch AI2026-08-18
OTIS Mock AIME 2024-202546.7%noneEpoch AI2026-08-06
ProofBench50%#29 of 77, top 38%Epoch AI
ProofBench16%maxEpoch AI
LMArena Math1443LMArena2026-10-08
LMArena Math1439LMArena2026-10-08
LMArena Math1455#59 of 285, top 21%LMArena2026-10-08

Knowledge

DeepSeek V4 Pro Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond90.9%highEpoch AI2026-08-06
GPQA Diamond91.7%#24 of 186, top 13%maxEpoch AI2026-08-18
GPQA Diamond89.6%maxEpoch AI2026-06-16
GPQA Diamond73.2%noneEpoch AI2026-08-06
SimpleQA Verified47%maxEpoch AI2026-08-27
SimpleQA Verified52.9%#21 of 77, top 28%maxEpoch AI2026-08-27
Vectara Hallucination Rate (lower is better)8.6%#40 of 96, top 42%Vectara Hallucination Leaderboard
LMArena Expert1464#57 of 273, top 21%LMArena2026-10-08
LMArena Expert1458LMArena2026-10-08
LMArena Expert1449LMArena2026-10-08

Multilingual

DeepSeek V4 Pro Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1439#45 of 297, top 16%LMArena2026-10-08
LMArena Non-English1431LMArena2026-10-08
LMArena Non-English1431LMArena2026-10-08
LMArena Chinese1480LMArena2026-10-08
LMArena Chinese1486#60 of 285, top 22%LMArena2026-10-08
LMArena Chinese1477LMArena2026-10-08
LMArena French1452LMArena2026-10-08
LMArena French1472#35 of 223, top 16%LMArena2026-10-08
LMArena French1448LMArena2026-10-08
LMArena German1458#37 of 231, top 17%LMArena2026-10-08
LMArena German1447LMArena2026-10-08
LMArena Japanese1415LMArena2026-10-08
LMArena Japanese1445#27 of 211, top 13%LMArena2026-10-08
LMArena Korean1411LMArena2026-10-08
LMArena Korean1425LMArena2026-10-08
LMArena Korean1447#20 of 213, top 10%LMArena2026-10-08
LMArena Russian1453#40 of 283, top 15%LMArena2026-10-08
LMArena Russian1448LMArena2026-10-08
LMArena Russian1449LMArena2026-10-08
LMArena Spanish1458#38 of 226, top 17%LMArena2026-10-08
LMArena Spanish1432LMArena2026-10-08
LMArena Spanish1449LMArena2026-10-08

Instruction Following

DeepSeek V4 Pro Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1445LMArena2026-10-08
LMArena Instruction Following1448#43 of 298, top 15%LMArena2026-10-08
LMArena Instruction Following1436LMArena2026-10-08

Long Context

DeepSeek V4 Pro Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
CL-bench Life13.5%#6 of 13, top 47%highEpoch AI
LMArena Longer Query1458#42 of 291, top 15%LMArena2026-10-08
LMArena Longer Query1446LMArena2026-10-08
LMArena Longer Query1458LMArena2026-10-08

Writing & Preference

DeepSeek V4 Pro Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1445LMArena2026-10-08
LMArena Text1446LMArena2026-10-08
LMArena Text1451#41 of 297, top 14%LMArena2026-10-08
LMArena Creative Writing1446#31 of 295, top 11%LMArena2026-10-08
LMArena Creative Writing1441LMArena2026-10-08
LMArena Creative Writing1438LMArena2026-10-08
EQ-Bench Creative Writing1553#50 of 115, top 44%EQ-Bench
EQ-Bench 41166#18 of 28, top 65%EQ-Bench
LMArena Multi-Turn1455LMArena2026-10-08
LMArena Multi-Turn1446LMArena2026-10-08
LMArena Multi-Turn1467#32 of 295, top 11%LMArena2026-10-08

API pricing by provider

DeepSeek V4 Pro API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$1.74$3.48—2026-10-10
deepinfra$1.30$2.60$0.102026-10-10
deepseek$0.66$1.98$0.0222026-10-10
openrouter$0.21$0.42$0.01742026-10-10
together$1.32$3.96$0.132026-10-10

Compare DeepSeek V4 Pro

Other DeepSeek models

Frequently asked questions

How good is DeepSeek V4 Pro?

DeepSeek V4 Pro by DeepSeek ranks 31st of 354 ranked models on the Noometry Index as of October 2026, with a score of 54.3. Its strongest category is reasoning, where it ranks 24th. API pricing starts at $0.66 per million input tokens and $1.98 per million output tokens, with a 1M-token context window.

How much does DeepSeek V4 Pro cost?

DeepSeek V4 Pro costs $0.66 per million input tokens and $1.98 per million output tokens on DeepSeek's own API, with cached input at $0.022.

What is DeepSeek V4 Pro's context window?

DeepSeek V4 Pro accepts up to 1M tokens of input and can write up to 393K tokens in one response.

Is DeepSeek V4 Pro open source?

Yes. DeepSeek V4 Pro's weights are downloadable from Hugging Face (deepseek-ai/DeepSeek-V4-Pro); check the license for commercial terms.

How fast is DeepSeek V4 Pro?

DeepSeek V4 Pro generated about 16 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are DeepSeek V4 Pro's strengths and weaknesses?

Relative to other ranked models, DeepSeek V4 Pro places best in reasoning, math, knowledge and lowest in agentic & tool use, long context, instruction following.

What is DeepSeek V4 Pro best at?

Its best category is reasoning, where it ranks 24th on Noometry.