DeepSeek, open weights
DeepSeek V4 Flash
DeepSeek V4 Flash by DeepSeek ranks 35th of 354 ranked models on the Noometry Index as of October 2026, with a score of 53.6. Its strongest category is reasoning, where it ranks 30th. API pricing starts at $0.15 per million input tokens and $0.60 per million output tokens, with a 1M-token context window.
Last verified
Specifications
- Noometry rank
- #35 of 354
- Index score
- 53.6
- Evidence
- Confirmed 41 results
- Provider
DeepSeek
- Released
- April 24, 2026
- Weights
- Open weights
- Reasoning
- Yes
- Context window
- 1M
- Max output
- 393K
- Input price
- $0.15 / M
- Output price
- $0.60 / M
- Blended price
- $0.26 / M
- Output speed
- 6 tokens/s Kagi
- Value
- #41 of 219
- Knowledge cutoff
- May 2025
- Input
- text, image
- Hugging Face
- deepseek-ai/DeepSeek-V4-Flash
Category scores
Each category score combines every public result we have in that category.
- Coding 47.9
- Reasoning 53.7
- Math 60.3
- Knowledge 55.4
- Multilingual 53.0
- Instruction Following 74.9
- Long Context 43.8
- Writing & Preference 63.8
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 47.9 | #59 | 5 |
| Reasoning | 53.7 | #30 | 11 |
| Math | 60.3 | #37 | 6 |
| Knowledge | 55.4 | #48 | 3 |
| Multilingual | 53.0 | #72 | 1 |
| Instruction Following | 74.9 | #81 | 1 |
| Long Context | 43.8 | #85 | 1 |
| Writing & Preference | 63.8 | #61 | 4 |
Strengths and weaknesses
Categories where DeepSeek V4 Flash places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Long Context | 43.8 | +2.9 | #85 of 296, top 29% |
| Instruction Following | 74.9 | +3.6 | #81 of 305, top 27% |
| Multilingual | 53.0 | +5.6 | #72 of 297, top 25% |
Closest competitors
The models ranked just above and below DeepSeek V4 Flash. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| DeepSeek V4 Pro | #31 | 54.3 | $0.99 | 16 | Compare |
| Gemini 3.5 Flash | #32 | 54.2 | $3.38 | — | Compare |
| Gemini 3.6 Flash | #33 | 54.1 | $1.50 | — | Compare |
| GPT-5.2 | #34 | 54.1 | $4.81 | 15 | Compare |
| GPT-6 Luna | #36 | 53.3 | $0.20 | — | Compare |
| Grok 4.7 | #37 | 53.1 | $3 | — | Compare |
| DeepSeek V4.1 Flash | #38 | 52.8 | $0.26 | — | Compare |
| GPT-5.2 Pro | #39 | 52.3 | $57.75 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierCode | 18.8% | #32 of 37, top 87% | Epoch AI | ||
| LMArena WebDev | 1431 | LMArena | 2026-10-08 | ||
| LMArena WebDev | 1582 | #28 of 113, top 25% | high | LMArena | 2026-10-08 |
| SciCode | 42% | high | Epoch AI | ||
| SciCode | 44.9% | max | Epoch AI | ||
| SciCode | 49.9% | #44 of 121, top 37% | max | Epoch AI | |
| WeirdML | 57% | high | Epoch AI | ||
| WeirdML | 43.8% | high | Epoch AI | ||
| WeirdML | 63% | #25 of 119, top 22% | max | Epoch AI | |
| WeirdML | 45.6% | max | Epoch AI | ||
| LMArena Coding | 1440 | LMArena | 2026-10-08 | ||
| LMArena Coding | 1457 | #75 of 294, top 26% | LMArena | 2026-10-08 | |
| ALE-Bench | 678.2 | high | Epoch AI | ||
| ALE-Bench | 1,306 | #25 of 105, top 24% | max | Epoch AI | |
| ALE-Bench | 324.98 | none | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 56% | high | Epoch AI | ||
| ARC-AGI-2 | 46% | low | Epoch AI | ||
| ARC-AGI-2 | 61.4% | #25 of 83, top 31% | max | Epoch AI | |
| ARC-AGI-2 | 2.1% | none | Epoch AI | ||
| SimpleBench | 61.1% | #21 of 77, top 28% | Epoch AI | ||
| SimpleBench | 46.3% | Epoch AI | |||
| Kagi LLM Benchmark | 52.2% | #61 of 99, top 62% | Kagi LLM Benchmark | ||
| NYT Connections (extended) | 89.6% | #20 of 91, top 22% | Lech Mazur benchmarks | ||
| ARC-AGI-1 | 87% | high | Epoch AI | ||
| ARC-AGI-1 | 84% | low | Epoch AI | ||
| ARC-AGI-1 | 89% | #28 of 83, top 34% | max | Epoch AI | |
| ARC-AGI-1 | 11.8% | none | Epoch AI | ||
| CritPt | 3.4% | high | Epoch AI | ||
| CritPt | 7.1% | max | Epoch AI | ||
| CritPt | 16.6% | #34 of 134, top 26% | max | Epoch AI | |
| Chess Puzzles | 33% | #31 of 129, top 25% | max | Epoch AI | 2026-08-02 |
| LMArena Hard Prompts | 1431 | LMArena | 2026-10-08 | ||
| LMArena Hard Prompts | 1444 | #75 of 297, top 26% | LMArena | 2026-10-08 | |
| Mystery Game Puzzles | 34% | #20 of 74, top 28% | max | Epoch AI | 2026-08-05 |
| DTBench | 90.9% | #32 of 151, top 22% | max | Epoch AI | |
| DTBench | 86.4% | max | Epoch AI | ||
| LMCA | 35.9% | max | Epoch AI | ||
| LMCA | 41.7% | #42 of 125, top 34% | max | Epoch AI | |
| Epoch Capabilities Index | 146.07 | Epoch AI | 2026-04-24 | ||
| Epoch Capabilities Index | 154.49 | #34 of 213, top 16% | Epoch AI | 2026-07-31 |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierMath (Tiers 1-3) | 57.5% | #38 of 81, top 47% | max | Epoch AI | 2026-08-02 |
| FrontierMath Tier 4 | 24.4% | #38 of 63, top 61% | max | Epoch AI | 2026-08-02 |
| MathArena Final-Answer Competitions | 76.5% | #8 of 29, top 28% | max | MathArena | |
| OTIS Mock AIME 2024-2025 | 94.4% | #38 of 173, top 22% | max | Epoch AI | 2026-08-02 |
| ProofBench | 56% | #23 of 77, top 30% | Epoch AI | ||
| LMArena Math | 1425 | LMArena | 2026-10-08 | ||
| LMArena Math | 1427 | #92 of 285, top 33% | LMArena | 2026-10-08 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 91% | #28 of 186, top 16% | max | Epoch AI | 2026-08-02 |
| SimpleQA Verified | 33.6% | #53 of 77, top 69% | max | Epoch AI | 2026-08-27 |
| LMArena Expert | 1436 | LMArena | 2026-10-08 | ||
| LMArena Expert | 1441 | #82 of 273, top 31% | LMArena | 2026-10-08 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1410 | LMArena | 2026-10-08 | ||
| LMArena Non-English | 1420 | #72 of 297, top 25% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1464 | LMArena | 2026-10-08 | ||
| LMArena Chinese | 1468 | #76 of 285, top 27% | LMArena | 2026-10-08 | |
| LMArena French | 1437 | LMArena | 2026-10-08 | ||
| LMArena French | 1439 | #84 of 223, top 38% | LMArena | 2026-10-08 | |
| LMArena German | 1418 | #78 of 231, top 34% | LMArena | 2026-10-08 | |
| LMArena German | 1415 | LMArena | 2026-10-08 | ||
| LMArena Japanese | 1406 | #54 of 211, top 26% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1389 | LMArena | 2026-10-08 | ||
| LMArena Korean | 1384 | #75 of 213, top 36% | LMArena | 2026-10-08 | |
| LMArena Korean | 1357 | LMArena | 2026-10-08 | ||
| LMArena Russian | 1428 | #70 of 283, top 25% | LMArena | 2026-10-08 | |
| LMArena Russian | 1426 | LMArena | 2026-10-08 | ||
| LMArena Spanish | 1418 | LMArena | 2026-10-08 | ||
| LMArena Spanish | 1436 | #72 of 226, top 32% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1421 | #71 of 298, top 24% | LMArena | 2026-10-08 | |
| LMArena Instruction Following | 1417 | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1434 | #74 of 291, top 26% | LMArena | 2026-10-08 | |
| LMArena Longer Query | 1424 | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1424 | LMArena | 2026-10-08 | ||
| LMArena Text | 1432 | #76 of 297, top 26% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1399 | LMArena | 2026-10-08 | ||
| LMArena Creative Writing | 1403 | #72 of 295, top 25% | LMArena | 2026-10-08 | |
| EQ-Bench Creative Writing | 1559 | #48 of 115, top 42% | EQ-Bench | ||
| EQ-Bench Creative Writing | 1441 | EQ-Bench | |||
| LMArena Multi-Turn | 1431 | LMArena | 2026-10-08 | ||
| LMArena Multi-Turn | 1449 | #57 of 295, top 20% | LMArena | 2026-10-08 |
API pricing by provider
Compare DeepSeek V4 Flash
- DeepSeek V4 Flash vs GPT-5.2
- DeepSeek V4 Flash vs GPT-6 Luna
- DeepSeek V4 Flash vs Gemini 3.6 Flash
- DeepSeek V4 Flash vs Grok 4.7
- DeepSeek V4 Flash vs Gemini 3.5 Flash
- DeepSeek V4 Flash vs DeepSeek V4.1 Flash
- DeepSeek V4 Flash vs GPT-6 Astra
- DeepSeek V4 Flash vs Claude Fable 5.1
- DeepSeek V4 Flash vs Gemini 3.8 Flash
- DeepSeek V4 Flash vs Kimi K3
- DeepSeek V4 Flash vs Grok 4.6
- DeepSeek V4 Flash vs Qwen3.8 Max
- DeepSeek V4 Flash vs GLM-5.3
- DeepSeek V4 Flash vs Muse Spark 1.3
Other DeepSeek models
Frequently asked questions
How good is DeepSeek V4 Flash?
DeepSeek V4 Flash by DeepSeek ranks 35th of 354 ranked models on the Noometry Index as of October 2026, with a score of 53.6. Its strongest category is reasoning, where it ranks 30th. API pricing starts at $0.15 per million input tokens and $0.60 per million output tokens, with a 1M-token context window.
How much does DeepSeek V4 Flash cost?
DeepSeek V4 Flash costs $0.15 per million input tokens and $0.60 per million output tokens on DeepSeek's own API, with cached input at $0.003.
What is DeepSeek V4 Flash's context window?
DeepSeek V4 Flash accepts up to 1M tokens of input and can write up to 393K tokens in one response.
Is DeepSeek V4 Flash open source?
Yes. DeepSeek V4 Flash's weights are downloadable from Hugging Face (deepseek-ai/DeepSeek-V4-Flash); check the license for commercial terms.
How fast is DeepSeek V4 Flash?
DeepSeek V4 Flash generated about 6 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are DeepSeek V4 Flash's strengths and weaknesses?
Relative to other ranked models, DeepSeek V4 Flash places best in reasoning, math, knowledge and lowest in long context, instruction following, multilingual.
What is DeepSeek V4 Flash best at?
Its best category is reasoning, where it ranks 30th on Noometry.