Model comparison

DeepSeek V4.1 Flash vs GPT-4.5

DeepSeek V4.1 Flash is the stronger model overall, scoring 52.8 to 37.2 on the Noometry Index.

Last verified . 21 shared benchmarks.

DeepSeek V4.1 Flash DeepSeek

52.8

Rank #38 Confirmed

GPT-4.5 OpenAI

37.2

Rank #208 Confirmed

Summary

  • They share 21 benchmarks with published results for both. DeepSeek V4.1 Flash scores higher in 10 categories and GPT-4.5 in 0 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where DeepSeek V4.1 Flash leads 50.2 to 13.9.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 98.3% for DeepSeek V4.1 Flash and 37.8% for GPT-4.5.
  • DeepSeek V4.1 Flash has downloadable open weights; the other is API-only.

Side by side

DeepSeek V4.1 Flash and GPT-4.5 specifications
DeepSeek V4.1 FlashGPT-4.5
ProviderDeepSeekOpenAI
Noometry Index52.837.2
Released2026-09-092025-02-27
WeightsOpenProprietary
Context window1M—
Max output393K—
Input $ / M tokens$0.15—
Output $ / M tokens$0.60—
Results tracked3742

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 52.9 (#32), GPT-4.5: 42.2 (#109)

Coding benchmarks
BenchmarkDeepSeek V4.1 FlashGPT-4.5
LMArena Coding15061396
Aider Polyglot—44.9%
LMArena WebDev1619—
SciCode51.9%—
WeirdML—39.4%
LiveBench Coding—75.2%
ALE-Bench1,092—

Agentic & Tool Use DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 31.2 (#69), GPT-4.5: 27.9 (#97)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek V4.1 FlashGPT-4.5
APEX-Agents39.5%—
Cybench—17.5%
GDP.pdf19.8%—

Reasoning DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 50.2 (#36), GPT-4.5: 13.9 (#330)

Reasoning benchmarks
BenchmarkDeepSeek V4.1 FlashGPT-4.5
LMArena Hard Prompts14831403
Epoch Capabilities Index154.9136.74
ARC-AGI-2—0.8%
SimpleBench—34.5%
NYT Connections (extended)89.6%—
ARC-AGI-1—10.3%
CritPt14.3%—
EnigmaEval—3.2%
LiveBench Reasoning—71.1%
Mystery Game Puzzles43%—
DTBench89.9%—
LiveBench Data Analysis—64.3%
LMCA47%—
Surface Evolver Bench46.3%—
ForecastBench—61.7
LiveBench—69%

Math DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 66.7 (#25), GPT-4.5: 32.6 (#211)

Math benchmarks
BenchmarkDeepSeek V4.1 FlashGPT-4.5
OTIS Mock AIME 2024-202598.3%37.8%
LMArena Math14771412
FrontierMath (Tiers 1-3)67.4%—
FrontierMath Tier 426.8%—
ProofBench54%—
LiveBench Math—69.3%
MATH Level 5—78.6%

Knowledge DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 57.9 (#38), GPT-4.5: 32.5 (#211)

Knowledge benchmarks
BenchmarkDeepSeek V4.1 FlashGPT-4.5
GPQA Diamond89.8%68.7%
LMArena Expert15061394
Humanity's Last Exam—5.4%
Confabulations—13.6%

Multimodal DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 39.1 (#61), GPT-4.5: 37.6 (#71)

Multimodal benchmarks
BenchmarkDeepSeek V4.1 FlashGPT-4.5
LMArena Vision12771195
VPCT—45%
Furniture Assembly34.2%—

Multilingual DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 55.0 (#35), GPT-4.5: 52.5 (#83)

Multilingual benchmarks
BenchmarkDeepSeek V4.1 FlashGPT-4.5
LMArena Non-English14481413
LMArena Chinese14971421
LMArena French14521418
LMArena German14841457
LMArena Japanese14121416
LMArena Korean14521392
LMArena Russian14711419
LMArena Spanish1459—

Instruction Following DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 77.3 (#26), GPT-4.5: 72.6 (#134)

Instruction Following benchmarks
BenchmarkDeepSeek V4.1 FlashGPT-4.5
LMArena Instruction Following14741404
LiveBench Instruction Following—72.3%

Long Context DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 45.2 (#47), GPT-4.5: 40.4 (#155)

Long Context benchmarks
BenchmarkDeepSeek V4.1 FlashGPT-4.5
LMArena Longer Query14751406
Fiction.LiveBench—63.9%

Writing & Preference DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 65.4 (#48), GPT-4.5: 56.9 (#134)

Writing & Preference benchmarks
BenchmarkDeepSeek V4.1 FlashGPT-4.5
LMArena Text14621417
LMArena Creative Writing14351394
EQ-Bench Creative Writing15401258
LMArena Multi-Turn14571444
Short-Story Creative Writing—75.6%
LiveBench Language—61.5%

Frequently asked questions

Is DeepSeek V4.1 Flash better than GPT-4.5?

DeepSeek V4.1 Flash is the stronger model overall, scoring 52.8 to 37.2 on the Noometry Index.

Is DeepSeek V4.1 Flash or GPT-4.5 better for coding?

DeepSeek V4.1 Flash scores higher on coding benchmarks: 52.9 versus 42.2 in the Noometry coding category.

How many benchmarks do DeepSeek V4.1 Flash and GPT-4.5 share?

21 benchmarks have published results for both models. DeepSeek V4.1 Flash has 37 scored results on Noometry and GPT-4.5 has 42.

Related comparisons

Go deeper