DeepSeek, open weights

DeepSeek-V3

DeepSeek-V3 by DeepSeek ranks 166th of 354 ranked models on the Noometry Index as of October 2026, with a score of 39.5. Its strongest category is coding, where it ranks 106th. API pricing starts at $0.24 per million input tokens and $0.90 per million output tokens, with a 164K-token context window.

Last verified

Specifications

Noometry rank
#166 of 354
Index score
39.5
Evidence
Confirmed 60 results
Provider
DeepSeek
Released
December 26, 2024
Weights
Open weights
Reasoning
No
Context window
164K
Max output
164K
Input price
$0.24 / M
Output price
$0.90 / M
Blended price
$0.41 / M
Output speed
73 tokens/s Kagi
Value
#62 of 219
Knowledge cutoff
July 2024
Input
text

Category scores

Each category score combines every public result we have in that category.

DeepSeek-V3 category scores
  1. Coding 42.3
  2. Reasoning 20.5
  3. Math 32.1
  4. Knowledge 37.5
  5. Multilingual 48.5
  6. Instruction Following 72.8
  7. Long Context 34.0
  8. Writing & Preference 57.4
DeepSeek-V3 category ranks
CategoryScoreRankResults
Coding42.3#1067
Reasoning20.5#2368
Math32.1#2195
Knowledge37.5#1556
Multilingual48.5#1431
Instruction Following72.8#1303
Long Context34.0#2532
Writing & Preference57.4#1307

Strengths and weaknesses

Categories where DeepSeek-V3 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

DeepSeek-V3: strongest categories
CategoryScorevs medianRank
Coding42.3+3.6#106 of 340, top 32%
Writing & Preference57.4+3.6#130 of 312, top 42%
Instruction Following72.8+1.5#130 of 305, top 43%

Weakest categories

DeepSeek-V3: weakest categories
CategoryScorevs medianRank
Long Context34.0−6.9#253 of 296, top 86%
Reasoning20.5−3.1#236 of 350, top 68%
Math32.1−4.5#219 of 327, top 67%

Closest competitors

The models ranked just above and below DeepSeek-V3. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to DeepSeek-V3
ModelRankScoreBlended $/MSpeed
DeepSeek-V3.2-Speciale#16239.7$0.85—Compare
Hunyuan Turbo 0110#16339.6——Compare
Claude 3.7 Sonnet#16439.5——Compare
Claude Haiku 4.5#16539.5$2—Compare
Grok 4 Fast#16739.4—577Compare
Olmo 3.1 32b Instruct#16839.4——Compare
Granite 4.2 3b#16939.4——Compare
Gemini 2.5 Flash#17039.3$0.85152Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

DeepSeek-V3 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
Aider Polyglot55.1%#18 of 44, top 41%Epoch AI
Aider Polyglot48.4%Epoch AI
SciCode35.4%Epoch AI
SciCode35.8%#95 of 121, top 79%Epoch AI
WeirdML36.1%#90 of 119, top 76%Epoch AI
BigCodeBench Instruct50%#2 of 64, top 4%BigCodeBench2024-12-26
LiveBench Coding70.9%#7 of 39, top 18%Epoch AI
LiveBench Coding61.8%Epoch AI
LMArena Coding1325LMArena2026-10-08
LMArena Coding1368#159 of 294, top 55%LMArena2026-10-08
BigCodeBench Complete62.2%Best of 66BigCodeBench2024-12-26
HumanEval+86.6%#5 of 45, top 12%nov 2024EvalPlus
MBPP+73%#10 of 38, top 27%nov 2024EvalPlus

Agentic & Tool Use

DeepSeek-V3 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
METR Time Horizons47.4%Epoch AI
METR Time Horizons49.6%#24 of 32, top 75%Epoch AI

Reasoning

DeepSeek-V3 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
SimpleBench18.9%Epoch AI
SimpleBench27.2%#62 of 77, top 81%Epoch AI
Kagi LLM Benchmark52.3%#58 of 99, top 59%Kagi LLM Benchmark
CritPt0%#103 of 134, top 77%Epoch AI
CritPt0%#103 of 134, top 77%Epoch AI
LiveBench Reasoning65.8%#12 of 39, top 31%Epoch AI
LiveBench Reasoning56.8%Epoch AI
LMArena Hard Prompts1312LMArena2026-10-08
LMArena Hard Prompts1365#150 of 297, top 51%LMArena2026-10-08
DTBench64.8%#102 of 151, top 68%Epoch AI
LiveBench Data Analysis60.4%Epoch AI
LiveBench Data Analysis60.9%#13 of 39, top 34%Epoch AI
LMCA15.5%#104 of 125, top 84%Epoch AI
BIG-Bench Hard87.5%#2 of 27, top 8%Epoch AI
Epoch Capabilities Index135.94#122 of 213, top 58%Epoch AI2025-03-24
Epoch Capabilities Index132.34Epoch AI2024-12-26
ForecastBench59.1#41 of 72, top 57%Epoch AI
HellaSwag88.9%#5 of 29, top 18%Epoch AI
LiveBench60.5%Epoch AI
LiveBench66.9%#10 of 39, top 26%Epoch AI
PIQA84.7%#6 of 27, top 23%Epoch AI
WinoGrande85.2%#7 of 43, top 17%Epoch AI

Math

DeepSeek-V3 Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-202537.8%#117 of 173, top 68%Epoch AI2025-04-01
OTIS Mock AIME 2024-202515.8%Epoch AI2025-02-25
Omni-MATH40.3%#26 of 57, top 46%HELM Capabilities
LiveBench Math73.5%#9 of 39, top 24%Epoch AI
LiveBench Math60.5%Epoch AI
LMArena Math1373#148 of 285, top 52%LMArena2026-10-08
LMArena Math1311LMArena2026-10-08
MATH Level 564.9%Epoch AI2025-01-27
MATH Level 575.5%#27 of 79, top 35%Epoch AI2025-04-01
FrontierMath (Feb 2025 set)1.7%#55 of 68, top 81%Epoch AI2025-03-07

Knowledge

DeepSeek-V3 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond67.6%#101 of 186, top 55%Epoch AI2025-04-01
GPQA Diamond56.5%Epoch AI2025-01-27
MMLU-Pro72.3%#32 of 58, top 56%HELM Capabilities
Confabulations (lower is better)26.1%#43 of 51, top 85%Lech Mazur benchmarks
Vectara Hallucination Rate (lower is better)6.1%#23 of 96, top 24%Vectara Hallucination Leaderboard
GPQA (HELM)53.8%#28 of 57, top 50%HELM Capabilities
LMArena Expert1306LMArena2026-10-08
LMArena Expert1351#156 of 273, top 58%LMArena2026-10-08
ARC (AI2) Challenge95.3%Best of 39Epoch AI
MMLU87.2%#3 of 81, top 4%Epoch AI
TriviaQA82.9%#7 of 25, top 29%Epoch AI

Multilingual

DeepSeek-V3 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1358#143 of 297, top 49%LMArena2026-10-08
LMArena Non-English1316LMArena2026-10-08
LMArena Chinese1338LMArena2026-10-08
LMArena Chinese1391#143 of 285, top 51%LMArena2026-10-08
LMArena French1343LMArena2026-10-08
LMArena French1385#129 of 223, top 58%LMArena2026-10-08
LMArena German1324LMArena2026-10-08
LMArena German1374#114 of 231, top 50%LMArena2026-10-08
LMArena Japanese1266LMArena2026-10-08
LMArena Japanese1333#108 of 211, top 52%LMArena2026-10-08
LMArena Korean1319#120 of 213, top 57%LMArena2026-10-08
LMArena Korean1248LMArena2026-10-08
LMArena Russian1373#133 of 283, top 47%LMArena2026-10-08
LMArena Russian1324LMArena2026-10-08
LMArena Spanish1358#137 of 226, top 61%LMArena2026-10-08
LMArena Spanish1352LMArena2026-10-08

Instruction Following

DeepSeek-V3 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LiveBench Instruction Following75.3%Epoch AI
LiveBench Instruction Following81.5%#8 of 39, top 21%Epoch AI
IFEval83.2%#30 of 57, top 53%HELM Capabilities
LMArena Instruction Following1315LMArena2026-10-08
LMArena Instruction Following1345#149 of 298, top 50%LMArena2026-10-08

Long Context

DeepSeek-V3 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
Fiction.LiveBench50%#34 of 47, top 73%Epoch AI
LMArena Longer Query1352#152 of 291, top 53%LMArena2026-10-08
LMArena Longer Query1343LMArena2026-10-08

Writing & Preference

DeepSeek-V3 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1333LMArena2026-10-08
LMArena Text1375#141 of 297, top 48%LMArena2026-10-08
LMArena Creative Writing1329LMArena2026-10-08
LMArena Creative Writing1364#116 of 295, top 40%LMArena2026-10-08
Short-Story Creative Writing77%#19 of 39, top 49%Epoch AI
EQ-Bench Creative Writing1472#62 of 115, top 54%EQ-Bench
WildBench83%#18 of 57, top 32%HELM Capabilities
LMArena Multi-Turn1389#130 of 295, top 45%LMArena2026-10-08
LMArena Multi-Turn1349LMArena2026-10-08
LiveBench Language49.1%#12 of 39, top 31%Epoch AI
LiveBench Language47.5%Epoch AI

API pricing by provider

DeepSeek-V3 API prices
RouteInput $/MOutput $/MCached input $/MChecked
deepinfra$0.24$0.90$0.142026-10-10
openrouter$0.32$0.89—2026-10-10
together$1.25$1.25—2026-10-10

Compare DeepSeek-V3

Other DeepSeek models

Frequently asked questions

How good is DeepSeek-V3?

DeepSeek-V3 by DeepSeek ranks 166th of 354 ranked models on the Noometry Index as of October 2026, with a score of 39.5. Its strongest category is coding, where it ranks 106th. API pricing starts at $0.24 per million input tokens and $0.90 per million output tokens, with a 164K-token context window.

How much does DeepSeek-V3 cost?

DeepSeek-V3 costs $0.24 per million input tokens and $0.90 per million output tokens on deepinfra, with cached input at $0.14.

What is DeepSeek-V3's context window?

DeepSeek-V3 accepts up to 164K tokens of input and can write up to 164K tokens in one response.

Is DeepSeek-V3 open source?

Yes. DeepSeek-V3's weights are downloadable from Hugging Face (deepseek-ai/DeepSeek-V3); check the license for commercial terms.

How fast is DeepSeek-V3?

DeepSeek-V3 generated about 73 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are DeepSeek-V3's strengths and weaknesses?

Relative to other ranked models, DeepSeek-V3 places best in coding, writing & preference, instruction following and lowest in long context, reasoning, math.

What is DeepSeek-V3 best at?

Its best category is coding, where it ranks 106th on Noometry.