DeepSeek, open weights
DeepSeek V4 Pro
DeepSeek V4 Pro by DeepSeek ranks 31st of 354 ranked models on the Noometry Index as of October 2026, with a score of 54.3. Its strongest category is reasoning, where it ranks 24th. API pricing starts at $0.66 per million input tokens and $1.98 per million output tokens, with a 1M-token context window.
Last verified
Specifications
- Noometry rank
- #31 of 354
- Index score
- 54.3
- Evidence
- Confirmed 48 results
- Provider
DeepSeek
- Released
- April 24, 2026
- Weights
- Open weights
- Reasoning
- Yes
- Context window
- 1M
- Max output
- 393K
- Input price
- $0.66 / M
- Output price
- $1.98 / M
- Blended price
- $0.99 / M
- Output speed
- 16 tokens/s Kagi
- Value
- #92 of 219
- Knowledge cutoff
- May 2025
- Input
- text
- Hugging Face
- deepseek-ai/DeepSeek-V4-Pro
Category scores
Each category score combines every public result we have in that category.
- Coding 52.4
- Agentic & Tool Use 32.8
- Reasoning 56.5
- Math 64.8
- Knowledge 59.5
- Multilingual 54.4
- Instruction Following 76.1
- Long Context 45.0
- Writing & Preference 65.5
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 52.4 | #34 | 6 |
| Agentic & Tool Use | 32.8 | #58 | 1 |
| Reasoning | 56.5 | #24 | 11 |
| Math | 64.8 | #30 | 6 |
| Knowledge | 59.5 | #31 | 4 |
| Multilingual | 54.4 | #45 | 1 |
| Instruction Following | 76.1 | #47 | 1 |
| Long Context | 45.0 | #51 | 2 |
| Writing & Preference | 65.5 | #46 | 5 |
Strengths and weaknesses
Categories where DeepSeek V4 Pro places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Agentic & Tool Use | 32.8 | +2.5 | #58 of 154, top 38% |
| Long Context | 45.0 | +4.0 | #51 of 296, top 18% |
| Instruction Following | 76.1 | +4.9 | #47 of 305, top 16% |
Closest competitors
The models ranked just above and below DeepSeek V4 Pro. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Muse Spark 1.3 | #27 | 54.8 | $2 | — | Compare |
| Gemini 3 Pro | #28 | 54.8 | — | 1 | Compare |
| Claude Sonnet 5 | #29 | 54.6 | $4 | — | Compare |
| GPT-5.6 Luna | #30 | 54.6 | $0.45 | 12 | Compare |
| Gemini 3.5 Flash | #32 | 54.2 | $3.38 | — | Compare |
| Gemini 3.6 Flash | #33 | 54.1 | $1.50 | — | Compare |
| GPT-5.2 | #34 | 54.1 | $4.81 | 15 | Compare |
| DeepSeek V4 Flash | #35 | 53.6 | $0.26 | 6 | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SWE-bench Verified | 77.6% | #6 of 32, top 19% | max | Epoch AI | 2026-06-18 |
| FrontierCode | 28.6% | #27 of 37, top 73% | high | Epoch AI | |
| FrontierCode | 17.6% | none | Epoch AI | ||
| LMArena WebDev | 1464 | LMArena | 2026-10-08 | ||
| LMArena WebDev | 1582 | #27 of 113, top 24% | LMArena | 2026-10-08 | |
| LMArena WebDev | 1447 | LMArena | 2026-10-08 | ||
| SciCode | 46.4% | high | Epoch AI | ||
| SciCode | 50% | max | Epoch AI | ||
| SciCode | 51% | #40 of 121, top 34% | max | Epoch AI | |
| SciCode | 39.9% | none | Epoch AI | ||
| WeirdML | 46.5% | high | Epoch AI | ||
| WeirdML | 66.2% | #22 of 119, top 19% | max | Epoch AI | |
| WeirdML | 48.9% | max | Epoch AI | ||
| LMArena Coding | 1469 | LMArena | 2026-10-08 | ||
| LMArena Coding | 1470 | #57 of 294, top 20% | LMArena | 2026-10-08 | |
| LMArena Coding | 1453 | LMArena | 2026-10-08 | ||
| ALE-Bench | 1,006 | high | Epoch AI | ||
| ALE-Bench | 1,403 | #19 of 105, top 19% | max | Epoch AI | |
| ALE-Bench | 521.67 | none | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| APEX-Agents | 47.3% | #29 of 49, top 60% | Epoch AI | ||
| Vending-Bench 2 | 3,285 | #41 of 60, top 69% | Epoch AI |
Reasoning
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierMath (Tiers 1-3) | 45.3% | max | Epoch AI | 2026-06-17 | |
| FrontierMath (Tiers 1-3) | 64.6% | #32 of 81, top 40% | max | Epoch AI | 2026-08-19 |
| FrontierMath Tier 4 | 2.4% | max | Epoch AI | 2026-06-17 | |
| FrontierMath Tier 4 | 26.8% | #34 of 63, top 54% | max | Epoch AI | 2026-08-19 |
| MathArena Final-Answer Competitions | 76.6% | #7 of 29, top 25% | max | MathArena | |
| OTIS Mock AIME 2024-2025 | 95.6% | high | Epoch AI | 2026-08-06 | |
| OTIS Mock AIME 2024-2025 | 96.7% | max | Epoch AI | 2026-06-17 | |
| OTIS Mock AIME 2024-2025 | 98.6% | #19 of 173, top 11% | max | Epoch AI | 2026-08-18 |
| OTIS Mock AIME 2024-2025 | 46.7% | none | Epoch AI | 2026-08-06 | |
| ProofBench | 50% | #29 of 77, top 38% | Epoch AI | ||
| ProofBench | 16% | max | Epoch AI | ||
| LMArena Math | 1443 | LMArena | 2026-10-08 | ||
| LMArena Math | 1439 | LMArena | 2026-10-08 | ||
| LMArena Math | 1455 | #59 of 285, top 21% | LMArena | 2026-10-08 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 90.9% | high | Epoch AI | 2026-08-06 | |
| GPQA Diamond | 91.7% | #24 of 186, top 13% | max | Epoch AI | 2026-08-18 |
| GPQA Diamond | 89.6% | max | Epoch AI | 2026-06-16 | |
| GPQA Diamond | 73.2% | none | Epoch AI | 2026-08-06 | |
| SimpleQA Verified | 47% | max | Epoch AI | 2026-08-27 | |
| SimpleQA Verified | 52.9% | #21 of 77, top 28% | max | Epoch AI | 2026-08-27 |
| Vectara Hallucination Rate (lower is better) | 8.6% | #40 of 96, top 42% | Vectara Hallucination Leaderboard | ||
| LMArena Expert | 1464 | #57 of 273, top 21% | LMArena | 2026-10-08 | |
| LMArena Expert | 1458 | LMArena | 2026-10-08 | ||
| LMArena Expert | 1449 | LMArena | 2026-10-08 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1439 | #45 of 297, top 16% | LMArena | 2026-10-08 | |
| LMArena Non-English | 1431 | LMArena | 2026-10-08 | ||
| LMArena Non-English | 1431 | LMArena | 2026-10-08 | ||
| LMArena Chinese | 1480 | LMArena | 2026-10-08 | ||
| LMArena Chinese | 1486 | #60 of 285, top 22% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1477 | LMArena | 2026-10-08 | ||
| LMArena French | 1452 | LMArena | 2026-10-08 | ||
| LMArena French | 1472 | #35 of 223, top 16% | LMArena | 2026-10-08 | |
| LMArena French | 1448 | LMArena | 2026-10-08 | ||
| LMArena German | 1458 | #37 of 231, top 17% | LMArena | 2026-10-08 | |
| LMArena German | 1447 | LMArena | 2026-10-08 | ||
| LMArena Japanese | 1415 | LMArena | 2026-10-08 | ||
| LMArena Japanese | 1445 | #27 of 211, top 13% | LMArena | 2026-10-08 | |
| LMArena Korean | 1411 | LMArena | 2026-10-08 | ||
| LMArena Korean | 1425 | LMArena | 2026-10-08 | ||
| LMArena Korean | 1447 | #20 of 213, top 10% | LMArena | 2026-10-08 | |
| LMArena Russian | 1453 | #40 of 283, top 15% | LMArena | 2026-10-08 | |
| LMArena Russian | 1448 | LMArena | 2026-10-08 | ||
| LMArena Russian | 1449 | LMArena | 2026-10-08 | ||
| LMArena Spanish | 1458 | #38 of 226, top 17% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1432 | LMArena | 2026-10-08 | ||
| LMArena Spanish | 1449 | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1445 | LMArena | 2026-10-08 | ||
| LMArena Instruction Following | 1448 | #43 of 298, top 15% | LMArena | 2026-10-08 | |
| LMArena Instruction Following | 1436 | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| CL-bench Life | 13.5% | #6 of 13, top 47% | high | Epoch AI | |
| LMArena Longer Query | 1458 | #42 of 291, top 15% | LMArena | 2026-10-08 | |
| LMArena Longer Query | 1446 | LMArena | 2026-10-08 | ||
| LMArena Longer Query | 1458 | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1445 | LMArena | 2026-10-08 | ||
| LMArena Text | 1446 | LMArena | 2026-10-08 | ||
| LMArena Text | 1451 | #41 of 297, top 14% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1446 | #31 of 295, top 11% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1441 | LMArena | 2026-10-08 | ||
| LMArena Creative Writing | 1438 | LMArena | 2026-10-08 | ||
| EQ-Bench Creative Writing | 1553 | #50 of 115, top 44% | EQ-Bench | ||
| EQ-Bench 4 | 1166 | #18 of 28, top 65% | EQ-Bench | ||
| LMArena Multi-Turn | 1455 | LMArena | 2026-10-08 | ||
| LMArena Multi-Turn | 1446 | LMArena | 2026-10-08 | ||
| LMArena Multi-Turn | 1467 | #32 of 295, top 11% | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| azure | $1.74 | $3.48 | — | 2026-10-10 |
| deepinfra | $1.30 | $2.60 | $0.10 | 2026-10-10 |
| deepseek | $0.66 | $1.98 | $0.022 | 2026-10-10 |
| openrouter | $0.21 | $0.42 | $0.0174 | 2026-10-10 |
| together | $1.32 | $3.96 | $0.13 | 2026-10-10 |
Compare DeepSeek V4 Pro
- DeepSeek V4 Pro vs GPT-5.6 Luna
- DeepSeek V4 Pro vs Gemini 3.5 Flash
- DeepSeek V4 Pro vs Claude Sonnet 5
- DeepSeek V4 Pro vs Gemini 3.6 Flash
- DeepSeek V4 Pro vs Gemini 3 Pro
- DeepSeek V4 Pro vs GPT-5.2
- DeepSeek V4 Pro vs GPT-6 Astra
- DeepSeek V4 Pro vs Claude Fable 5.1
- DeepSeek V4 Pro vs Gemini 3.8 Flash
- DeepSeek V4 Pro vs Kimi K3
- DeepSeek V4 Pro vs Grok 4.6
- DeepSeek V4 Pro vs Qwen3.8 Max
- DeepSeek V4 Pro vs GLM-5.3
- DeepSeek V4 Pro vs Muse Spark 1.3
Other DeepSeek models
Frequently asked questions
How good is DeepSeek V4 Pro?
DeepSeek V4 Pro by DeepSeek ranks 31st of 354 ranked models on the Noometry Index as of October 2026, with a score of 54.3. Its strongest category is reasoning, where it ranks 24th. API pricing starts at $0.66 per million input tokens and $1.98 per million output tokens, with a 1M-token context window.
How much does DeepSeek V4 Pro cost?
DeepSeek V4 Pro costs $0.66 per million input tokens and $1.98 per million output tokens on DeepSeek's own API, with cached input at $0.022.
What is DeepSeek V4 Pro's context window?
DeepSeek V4 Pro accepts up to 1M tokens of input and can write up to 393K tokens in one response.
Is DeepSeek V4 Pro open source?
Yes. DeepSeek V4 Pro's weights are downloadable from Hugging Face (deepseek-ai/DeepSeek-V4-Pro); check the license for commercial terms.
How fast is DeepSeek V4 Pro?
DeepSeek V4 Pro generated about 16 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are DeepSeek V4 Pro's strengths and weaknesses?
Relative to other ranked models, DeepSeek V4 Pro places best in reasoning, math, knowledge and lowest in agentic & tool use, long context, instruction following.
What is DeepSeek V4 Pro best at?
Its best category is reasoning, where it ranks 24th on Noometry.