Model comparison

DeepSeek-V3 vs GPT-4o

DeepSeek-V3 is the stronger model overall, scoring 39.5 to 28.6 on the Noometry Index.

Last verified . 51 shared benchmarks.

DeepSeek-V3 DeepSeek

39.5

Rank #166 Confirmed

GPT-4o OpenAI

28.6

Rank #324 Confirmed

Summary

  • They share 51 benchmarks with published results for both. DeepSeek-V3 scores higher in 7 categories and GPT-4o in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where DeepSeek-V3 leads 32.1 to 10.6.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 37.8% for DeepSeek-V3 and 6.4% for GPT-4o.
  • DeepSeek-V3 is cheaper at $0.24 / $0.90 per million input/output tokens, against $2.50 / $10 for GPT-4o.
  • DeepSeek-V3 accepts more context: 164K tokens versus 128K.
  • DeepSeek-V3 has downloadable open weights; the other is API-only.

Side by side

DeepSeek-V3 and GPT-4o specifications
DeepSeek-V3GPT-4o
ProviderDeepSeekOpenAI
Noometry Index39.528.6
Released2024-12-262024-05-13
WeightsOpenProprietary
Context window164K128K
Max output164K16K
Input $ / M tokens$0.24$2.50
Output $ / M tokens$0.90$10
Results tracked6072

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3 leads

DeepSeek-V3: 42.3 (#106), GPT-4o: 24.8 (#328)

Coding benchmarks
BenchmarkDeepSeek-V3GPT-4o
Aider Polyglot55.1%45.3%
WeirdML36.1%25.1%
BigCodeBench Instruct50%51.1%
LiveBench Coding70.9%51.4%
LMArena Coding13681297
BigCodeBench Complete62.2%61.1%
HumanEval+86.6%87.2%
MBPP+73%72.2%
SWE-bench Verified—31%
SWE-bench Verified (bash only)—21.6%
SciCode35.8%—
GSO—0%
CadEval—26%

Agentic & Tool Use Not comparable

DeepSeek-V3: —, GPT-4o: 21.0 (#141)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3GPT-4o
METR Time Horizons49.6%40.8%
GDPval—9.9%
TheAgentCompany—8.6%
Cybench—12.5%
BALROG—32.3%
LMArena Search—1006

Reasoning DeepSeek-V3 leads

DeepSeek-V3: 20.5 (#236), GPT-4o: 9.4 (#343)

Reasoning benchmarks
BenchmarkDeepSeek-V3GPT-4o
SimpleBench27.2%17.8%
CritPt0%0%
LiveBench Reasoning65.8%55.8%
LMArena Hard Prompts13651281
DTBench64.8%64.5%
LiveBench Data Analysis60.9%60.9%
LMCA15.5%16.6%
Epoch Capabilities Index135.94128.97
ForecastBench59.157.7
LiveBench66.9%55.3%
ARC-AGI-2—0%
Kagi LLM Benchmark52.3%—
ARC-AGI-1—4.5%
Chess Puzzles—13%
EnigmaEval—0.8%
BIG-Bench Hard87.5%—
HellaSwag88.9%—
PIQA84.7%—
WinoGrande85.2%—

Math DeepSeek-V3 leads

DeepSeek-V3: 32.1 (#219), GPT-4o: 10.6 (#312)

Math benchmarks
BenchmarkDeepSeek-V3GPT-4o
OTIS Mock AIME 2024-202537.8%6.4%
Omni-MATH40.3%29.3%
LiveBench Math73.5%49.5%
LMArena Math13731285
MATH Level 575.5%53.3%
FrontierMath (Feb 2025 set)1.7%0.3%
FrontierMath (Tiers 1-3)—0.4%

Knowledge DeepSeek-V3 leads

DeepSeek-V3: 37.5 (#155), GPT-4o: 28.8 (#242)

Knowledge benchmarks
BenchmarkDeepSeek-V3GPT-4o
GPQA Diamond67.6%49.2%
MMLU-Pro72.3%71.3%
Confabulations26.1%15.3%
Vectara Hallucination Rate6.1%9.6%
GPQA (HELM)53.8%52%
LMArena Expert13511250
MMLU87.2%88.1%
Humanity's Last Exam—2.7%
SimpleQA Verified—26%
ARC (AI2) Challenge95.3%—
TriviaQA82.9%—

Multimodal Not comparable

DeepSeek-V3: —, GPT-4o: 34.5 (#91)

Multimodal benchmarks
BenchmarkDeepSeek-V3GPT-4o
LMArena Vision—1137
Video-MME—71.9%
GeoBench—71%
VPCT—40%
ScienceQA—88.5%

Multilingual DeepSeek-V3 leads

DeepSeek-V3: 48.5 (#143), GPT-4o: 43.2 (#186)

Multilingual benchmarks
BenchmarkDeepSeek-V3GPT-4o
LMArena Non-English13581283
LMArena Chinese13911277
LMArena French13851304
LMArena German13741282
LMArena Japanese13331257
LMArena Korean13191234
LMArena Russian13731286
LMArena Spanish13581292

Instruction Following DeepSeek-V3 leads

DeepSeek-V3: 72.8 (#130), GPT-4o: 66.6 (#207)

Instruction Following benchmarks
BenchmarkDeepSeek-V3GPT-4o
LiveBench Instruction Following81.5%68.6%
IFEval83.2%81.7%
LMArena Instruction Following13451278

Long Context GPT-4o leads

DeepSeek-V3: 34.0 (#253), GPT-4o: 39.4 (#179)

Long Context benchmarks
BenchmarkDeepSeek-V3GPT-4o
Fiction.LiveBench50%66.7%
LMArena Longer Query13521289

Writing & Preference DeepSeek-V3 leads

DeepSeek-V3: 57.4 (#130), GPT-4o: 52.6 (#166)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3GPT-4o
LMArena Text13751300
LMArena Creative Writing13641292
Short-Story Creative Writing77%81.8%
WildBench83%82.8%
LMArena Multi-Turn13891302
LiveBench Language49.1%47.6%
EQ-Bench Creative Writing1472—

Frequently asked questions

Is DeepSeek-V3 better than GPT-4o?

DeepSeek-V3 is the stronger model overall, scoring 39.5 to 28.6 on the Noometry Index.

Which is cheaper, DeepSeek-V3 or GPT-4o?

DeepSeek-V3 is cheaper. It lists at $0.24 per million input tokens and $0.90 per million output tokens; GPT-4o lists at $2.50 and $10.

Is DeepSeek-V3 or GPT-4o better for coding?

DeepSeek-V3 scores higher on coding benchmarks: 42.3 versus 24.8 in the Noometry coding category.

Which has the bigger context window?

DeepSeek-V3 does, with 164K tokens against 128K.

How many benchmarks do DeepSeek-V3 and GPT-4o share?

51 benchmarks have published results for both models. DeepSeek-V3 has 60 scored results on Noometry and GPT-4o has 72.

Related comparisons

Go deeper