Model comparison

DeepSeek V4 Pro vs GPT-5 Mini

DeepSeek V4 Pro is the stronger model overall, scoring 54.3 to 41.8 on the Noometry Index.

Last verified . 40 shared benchmarks.

DeepSeek V4 Pro DeepSeek

54.3

Rank #31 Confirmed

GPT-5 Mini OpenAI

41.8

Rank #128 Confirmed

Summary

  • They share 40 benchmarks with published results for both. DeepSeek V4 Pro scores higher in 8 categories and GPT-5 Mini in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where DeepSeek V4 Pro leads 56.5 to 23.9.
  • The biggest single-benchmark swing is ARC-AGI-2: 61.3% for DeepSeek V4 Pro and 4.4% for GPT-5 Mini.
  • GPT-5 Mini is cheaper at $0.25 / $2 per million input/output tokens, against $0.66 / $1.98 for DeepSeek V4 Pro.
  • DeepSeek V4 Pro accepts more context: 1M tokens versus 400K.
  • DeepSeek V4 Pro has downloadable open weights; the other is API-only.

Side by side

DeepSeek V4 Pro and GPT-5 Mini specifications
DeepSeek V4 ProGPT-5 Mini
ProviderDeepSeekOpenAI
Noometry Index54.341.8
Released2026-04-242025-08-07
WeightsOpenProprietary
Context window1M400K
Max output393K128K
Input $ / M tokens$0.66$0.25
Output $ / M tokens$1.98$2
Results tracked4860

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek V4 Pro leads

DeepSeek V4 Pro: 52.4 (#34), GPT-5 Mini: 40.1 (#146)

Coding benchmarks
BenchmarkDeepSeek V4 ProGPT-5 Mini
SWE-bench Verified77.6%64.7%
SciCode51%39.2%
WeirdML66.2%52.7%
LMArena Coding14701406
ALE-Bench1,403799.77
FrontierCode28.6%—
SWE-bench Verified (bash only)—59.8%
LMArena WebDev1582—
SWE-bench Multilingual—39.7%
AlgoTune—1.38

Agentic & Tool Use DeepSeek V4 Pro leads

DeepSeek V4 Pro: 32.8 (#58), GPT-5 Mini: 31.1 (#70)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek V4 ProGPT-5 Mini
Vending-Bench 23,285-31.18
Terminal-Bench—34.8%
APEX-Agents47.3%—
Berkeley Function Calling Leaderboard—55.5%

Reasoning DeepSeek V4 Pro leads

DeepSeek V4 Pro: 56.5 (#24), GPT-5 Mini: 23.9 (#168)

Reasoning benchmarks
BenchmarkDeepSeek V4 ProGPT-5 Mini
ARC-AGI-261.3%4.4%
Kagi LLM Benchmark53.5%70.3%
ARC-AGI-190.5%54.3%
CritPt18%0%
Chess Puzzles47%30%
LMArena Hard Prompts14611380
Mystery Game Puzzles43%10%
DTBench93.9%80.5%
LMCA45.5%34.2%
Epoch Capabilities Index155.31145.52
ForecastBench56.161
NYT Connections (extended)91.3%—
EnigmaEval—8.2%
Surface Evolver Bench40%—

Math DeepSeek V4 Pro leads

DeepSeek V4 Pro: 64.8 (#30), GPT-5 Mini: 46.7 (#69)

Math benchmarks
BenchmarkDeepSeek V4 ProGPT-5 Mini
FrontierMath (Tiers 1-3)64.6%46.7%
FrontierMath Tier 426.8%12.2%
OTIS Mock AIME 2024-202598.6%86.7%
ProofBench50%9%
LMArena Math14551378
MathArena Final-Answer Competitions76.6%—
Omni-MATH—72.2%
MATH Level 5—97.8%
FrontierMath (Feb 2025 set)—27.2%
FrontierMath Tier 4 (v1)—6.3%

Knowledge DeepSeek V4 Pro leads

DeepSeek V4 Pro: 59.5 (#31), GPT-5 Mini: 45.6 (#86)

Knowledge benchmarks
BenchmarkDeepSeek V4 ProGPT-5 Mini
GPQA Diamond91.7%75%
SimpleQA Verified52.9%21.6%
Vectara Hallucination Rate8.6%12.9%
LMArena Expert14641379
Humanity's Last Exam—19.4%
MMLU-Pro—83.5%
Confabulations—13.3%
GPQA (HELM)—75.6%

Multimodal Not comparable

DeepSeek V4 Pro: —, GPT-5 Mini: 35.6 (#85)

Multimodal benchmarks
BenchmarkDeepSeek V4 ProGPT-5 Mini
LMArena Vision—1202
VPCT—40.2%

Multilingual DeepSeek V4 Pro leads

DeepSeek V4 Pro: 54.4 (#45), GPT-5 Mini: 48.9 (#137)

Multilingual benchmarks
BenchmarkDeepSeek V4 ProGPT-5 Mini
LMArena Non-English14391363
LMArena Chinese14861385
LMArena French14721386
LMArena German14581366
LMArena Japanese14451341
LMArena Korean14471308
LMArena Russian14531362
LMArena Spanish14581355

Instruction Following Too close to call

DeepSeek V4 Pro: 76.1 (#47), GPT-5 Mini: 76.2 (#46)

Instruction Following benchmarks
BenchmarkDeepSeek V4 ProGPT-5 Mini
LMArena Instruction Following14481357
IFEval—92.7%

Long Context DeepSeek V4 Pro leads

DeepSeek V4 Pro: 45.0 (#51), GPT-5 Mini: 41.9 (#132)

Long Context benchmarks
BenchmarkDeepSeek V4 ProGPT-5 Mini
LMArena Longer Query14581355
Fiction.LiveBench—69.4%
CL-bench Life13.5%—

Writing & Preference DeepSeek V4 Pro leads

DeepSeek V4 Pro: 65.5 (#46), GPT-5 Mini: 55.2 (#148)

Writing & Preference benchmarks
BenchmarkDeepSeek V4 ProGPT-5 Mini
LMArena Text14511373
LMArena Creative Writing14461325
EQ-Bench Creative Writing15531313
LMArena Multi-Turn14671363
Short-Story Creative Writing—83.1%
WildBench—85.5%
EQ-Bench 41166—

Frequently asked questions

Is DeepSeek V4 Pro better than GPT-5 Mini?

DeepSeek V4 Pro is the stronger model overall, scoring 54.3 to 41.8 on the Noometry Index.

Which is cheaper, DeepSeek V4 Pro or GPT-5 Mini?

GPT-5 Mini is cheaper. It lists at $0.25 per million input tokens and $2 per million output tokens; DeepSeek V4 Pro lists at $0.66 and $1.98.

Is DeepSeek V4 Pro or GPT-5 Mini better for coding?

DeepSeek V4 Pro scores higher on coding benchmarks: 52.4 versus 40.1 in the Noometry coding category.

Which has the bigger context window?

DeepSeek V4 Pro does, with 1M tokens against 400K.

How many benchmarks do DeepSeek V4 Pro and GPT-5 Mini share?

40 benchmarks have published results for both models. DeepSeek V4 Pro has 48 scored results on Noometry and GPT-5 Mini has 60.

Related comparisons

Go deeper