Model comparison

DeepSeek-V3.2-Exp vs GPT-4.1

DeepSeek-V3.2-Exp is the stronger model overall, scoring 44.3 to 35.9 on the Noometry Index.

Last verified . 35 shared benchmarks.

DeepSeek-V3.2-Exp DeepSeek

44.3

Rank #78 Confirmed

GPT-4.1 OpenAI

35.9

Rank #219 Confirmed

Summary

  • They share 35 benchmarks with published results for both. DeepSeek-V3.2-Exp scores higher in 8 categories and GPT-4.1 in 1 category; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where DeepSeek-V3.2-Exp leads 41.7 to 22.3.
  • The biggest single-benchmark swing is ARC-AGI-1: 57% for DeepSeek-V3.2-Exp and 5.5% for GPT-4.1.
  • DeepSeek-V3.2-Exp is cheaper at $0.26 / $0.38 per million input/output tokens, against $2 / $8 for GPT-4.1.
  • GPT-4.1 accepts more context: 1.05M tokens versus 164K.
  • DeepSeek-V3.2-Exp has downloadable open weights; the other is API-only.

Side by side

DeepSeek-V3.2-Exp and GPT-4.1 specifications
DeepSeek-V3.2-ExpGPT-4.1
ProviderDeepSeekOpenAI
Noometry Index44.335.9
Released2025-09-292025-04-14
WeightsOpenProprietary
Context window164K1.05M
Max output66K33K
Input $ / M tokens$0.26$2
Output $ / M tokens$0.38$8
Results tracked4952

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 46.5 (#65), GPT-4.1: 34.4 (#238)

Coding benchmarks
BenchmarkDeepSeek-V3.2-ExpGPT-4.1
SWE-bench Verified (bash only)70%39.6%
Aider Polyglot74.2%52.4%
WeirdML39.5%39%
LMArena Coding14541391
SWE-bench Verified—48.5%
LMArena WebDev1362—
SWE-bench Multilingual59%—
SciCode38.9%—
CadEval—42%
ALE-Bench—558.1

Agentic & Tool Use GPT-4.1 leads

DeepSeek-V3.2-Exp: 32.7 (#59), GPT-4.1: 34.7 (#43)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3.2-ExpGPT-4.1
Berkeley Function Calling Leaderboard56.7%54%
Terminal-Bench39.6%—
APEX-Agents21.3%—
TheAgentCompany42.9%—
Vending-Bench 21,034—

Reasoning DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 22.1 (#208), GPT-4.1: 11.7 (#339)

Reasoning benchmarks
BenchmarkDeepSeek-V3.2-ExpGPT-4.1
ARC-AGI-24%0.4%
Kagi LLM Benchmark52.2%52.3%
ARC-AGI-157%5.5%
Chess Puzzles14%6%
LMArena Hard Prompts14341384
DTBench87.7%68.3%
LMCA29.1%25.6%
Epoch Capabilities Index146.27136.78
SimpleBench—27%
NYT Connections (extended)36.7%—
CritPt2.9%—
EnigmaEval—2.2%
Thematic Generalization65%—
ForecastBench—61.5

Math DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 41.7 (#87), GPT-4.1: 22.3 (#280)

Knowledge DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 51.7 (#66), GPT-4.1: 37.1 (#160)

Knowledge benchmarks
BenchmarkDeepSeek-V3.2-ExpGPT-4.1
GPQA Diamond83.4%66.9%
Vectara Hallucination Rate5.3%5.6%
LMArena Expert14361364
Humanity's Last Exam—5.4%
SimpleQA Verified—31.1%
MMLU-Pro—81.1%
GPQA (HELM)—65.9%

Multimodal Not comparable

DeepSeek-V3.2-Exp: —, GPT-4.1: 38.2 (#67)

Multimodal benchmarks
BenchmarkDeepSeek-V3.2-ExpGPT-4.1
LMArena Vision—1211
GeoBench—72%

Multilingual DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 52.2 (#90), GPT-4.1: 49.4 (#133)

Multilingual benchmarks
BenchmarkDeepSeek-V3.2-ExpGPT-4.1
LMArena Non-English14091370
LMArena Chinese14611382
LMArena French14331382
LMArena German14401381
LMArena Japanese13741319
LMArena Korean13711339
LMArena Russian14241377
LMArena Spanish14401376

Instruction Following DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 74.5 (#93), GPT-4.1: 71.3 (#153)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.2-ExpGPT-4.1
LMArena Instruction Following14131367
IFEval—83.8%

Long Context DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 47.6 (#16), GPT-4.1: 40.0 (#163)

Long Context benchmarks
BenchmarkDeepSeek-V3.2-ExpGPT-4.1
Fiction.LiveBench83.3%63.9%
LMArena Longer Query14281385
CL-bench13.2%—
CL-bench Life9.5%—

Writing & Preference DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 62.4 (#77), GPT-4.1: 57.6 (#125)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.2-ExpGPT-4.1
LMArena Text14251383
LMArena Creative Writing14031363
EQ-Bench Creative Writing15151420
LMArena Multi-Turn14271398
WildBench—85.4%

Frequently asked questions

Is DeepSeek-V3.2-Exp better than GPT-4.1?

DeepSeek-V3.2-Exp is the stronger model overall, scoring 44.3 to 35.9 on the Noometry Index.

Which is cheaper, DeepSeek-V3.2-Exp or GPT-4.1?

DeepSeek-V3.2-Exp is cheaper. It lists at $0.26 per million input tokens and $0.38 per million output tokens; GPT-4.1 lists at $2 and $8.

Is DeepSeek-V3.2-Exp or GPT-4.1 better for coding?

DeepSeek-V3.2-Exp scores higher on coding benchmarks: 46.5 versus 34.4 in the Noometry coding category.

Which has the bigger context window?

GPT-4.1 does, with 1.05M tokens against 164K.

How many benchmarks do DeepSeek-V3.2-Exp and GPT-4.1 share?

35 benchmarks have published results for both models. DeepSeek-V3.2-Exp has 49 scored results on Noometry and GPT-4.1 has 52.

Related comparisons

Go deeper