Model comparison

DeepSeek V4 Pro vs GPT-4.1

DeepSeek V4 Pro is the stronger model overall, scoring 54.3 to 35.9 on the Noometry Index.

Last verified . 34 shared benchmarks.

DeepSeek V4 Pro DeepSeek

54.3

Rank #31 Confirmed

GPT-4.1 OpenAI

35.9

Rank #219 Confirmed

Summary

  • They share 34 benchmarks with published results for both. DeepSeek V4 Pro scores higher in 8 categories and GPT-4.1 in 1 category; 9 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where DeepSeek V4 Pro leads 56.5 to 11.7.
  • The biggest single-benchmark swing is ARC-AGI-1: 90.5% for DeepSeek V4 Pro and 5.5% for GPT-4.1.
  • DeepSeek V4 Pro is cheaper at $0.66 / $1.98 per million input/output tokens, against $2 / $8 for GPT-4.1.
  • GPT-4.1 accepts more context: 1.05M tokens versus 1M.
  • DeepSeek V4 Pro has downloadable open weights; the other is API-only.

Side by side

DeepSeek V4 Pro and GPT-4.1 specifications
DeepSeek V4 ProGPT-4.1
ProviderDeepSeekOpenAI
Noometry Index54.335.9
Released2026-04-242025-04-14
WeightsOpenProprietary
Context window1M1.05M
Max output393K33K
Input $ / M tokens$0.66$2
Output $ / M tokens$1.98$8
Results tracked4852

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek V4 Pro leads

DeepSeek V4 Pro: 52.4 (#34), GPT-4.1: 34.4 (#238)

Coding benchmarks
BenchmarkDeepSeek V4 ProGPT-4.1
SWE-bench Verified77.6%48.5%
WeirdML66.2%39%
LMArena Coding14701391
ALE-Bench1,403558.1
FrontierCode28.6%—
SWE-bench Verified (bash only)—39.6%
Aider Polyglot—52.4%
LMArena WebDev1582—
SciCode51%—
CadEval—42%

Agentic & Tool Use GPT-4.1 leads

DeepSeek V4 Pro: 32.8 (#58), GPT-4.1: 34.7 (#43)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek V4 ProGPT-4.1
APEX-Agents47.3%—
Berkeley Function Calling Leaderboard—54%
Vending-Bench 23,285—

Reasoning DeepSeek V4 Pro leads

DeepSeek V4 Pro: 56.5 (#24), GPT-4.1: 11.7 (#339)

Reasoning benchmarks
BenchmarkDeepSeek V4 ProGPT-4.1
ARC-AGI-261.3%0.4%
Kagi LLM Benchmark53.5%52.3%
ARC-AGI-190.5%5.5%
Chess Puzzles47%6%
LMArena Hard Prompts14611384
DTBench93.9%68.3%
LMCA45.5%25.6%
Epoch Capabilities Index155.31136.78
ForecastBench56.161.5
SimpleBench—27%
NYT Connections (extended)91.3%—
CritPt18%—
EnigmaEval—2.2%
Mystery Game Puzzles43%—
Surface Evolver Bench40%—

Math DeepSeek V4 Pro leads

DeepSeek V4 Pro: 64.8 (#30), GPT-4.1: 22.3 (#280)

Knowledge DeepSeek V4 Pro leads

DeepSeek V4 Pro: 59.5 (#31), GPT-4.1: 37.1 (#160)

Knowledge benchmarks
BenchmarkDeepSeek V4 ProGPT-4.1
GPQA Diamond91.7%66.9%
SimpleQA Verified52.9%31.1%
Vectara Hallucination Rate8.6%5.6%
LMArena Expert14641364
Humanity's Last Exam—5.4%
MMLU-Pro—81.1%
GPQA (HELM)—65.9%

Multimodal Not comparable

DeepSeek V4 Pro: —, GPT-4.1: 38.2 (#67)

Multimodal benchmarks
BenchmarkDeepSeek V4 ProGPT-4.1
LMArena Vision—1211
GeoBench—72%

Multilingual DeepSeek V4 Pro leads

DeepSeek V4 Pro: 54.4 (#45), GPT-4.1: 49.4 (#133)

Multilingual benchmarks
BenchmarkDeepSeek V4 ProGPT-4.1
LMArena Non-English14391370
LMArena Chinese14861382
LMArena French14721382
LMArena German14581381
LMArena Japanese14451319
LMArena Korean14471339
LMArena Russian14531377
LMArena Spanish14581376

Instruction Following DeepSeek V4 Pro leads

DeepSeek V4 Pro: 76.1 (#47), GPT-4.1: 71.3 (#153)

Instruction Following benchmarks
BenchmarkDeepSeek V4 ProGPT-4.1
LMArena Instruction Following14481367
IFEval—83.8%

Long Context DeepSeek V4 Pro leads

DeepSeek V4 Pro: 45.0 (#51), GPT-4.1: 40.0 (#163)

Long Context benchmarks
BenchmarkDeepSeek V4 ProGPT-4.1
LMArena Longer Query14581385
Fiction.LiveBench—63.9%
CL-bench Life13.5%—

Writing & Preference DeepSeek V4 Pro leads

DeepSeek V4 Pro: 65.5 (#46), GPT-4.1: 57.6 (#125)

Writing & Preference benchmarks
BenchmarkDeepSeek V4 ProGPT-4.1
LMArena Text14511383
LMArena Creative Writing14461363
EQ-Bench Creative Writing15531420
LMArena Multi-Turn14671398
WildBench—85.4%
EQ-Bench 41166—

Frequently asked questions

Is DeepSeek V4 Pro better than GPT-4.1?

DeepSeek V4 Pro is the stronger model overall, scoring 54.3 to 35.9 on the Noometry Index.

Which is cheaper, DeepSeek V4 Pro or GPT-4.1?

DeepSeek V4 Pro is cheaper. It lists at $0.66 per million input tokens and $1.98 per million output tokens; GPT-4.1 lists at $2 and $8.

Is DeepSeek V4 Pro or GPT-4.1 better for coding?

DeepSeek V4 Pro scores higher on coding benchmarks: 52.4 versus 34.4 in the Noometry coding category.

Which has the bigger context window?

GPT-4.1 does, with 1.05M tokens against 1M.

How many benchmarks do DeepSeek V4 Pro and GPT-4.1 share?

34 benchmarks have published results for both models. DeepSeek V4 Pro has 48 scored results on Noometry and GPT-4.1 has 52.

Related comparisons

Go deeper