Model comparison

DeepSeek-V3.2-Exp vs GPT-5.1

GPT-5.1 is the stronger model overall, scoring 49.0 to 44.3 on the Noometry Index. DeepSeek-V3.2-Exp costs 12× less per token, which makes it the better buy when GPT-5.1's lead doesn't matter for your workload.

Last verified . 37 shared benchmarks.

DeepSeek-V3.2-Exp DeepSeek

44.3

Rank #78 Confirmed

GPT-5.1 OpenAI

49.0

Rank #53 Confirmed

Summary

  • They share 37 benchmarks with published results for both. DeepSeek-V3.2-Exp scores higher in 3 categories and GPT-5.1 in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where GPT-5.1 leads 39.8 to 22.1.
  • The biggest single-benchmark swing is WeirdML: 39.5% for DeepSeek-V3.2-Exp and 60.8% for GPT-5.1.
  • DeepSeek-V3.2-Exp is cheaper at $0.26 / $0.38 per million input/output tokens, against $1.25 / $10 for GPT-5.1.
  • GPT-5.1 accepts more context: 400K tokens versus 164K.
  • DeepSeek-V3.2-Exp has downloadable open weights; the other is API-only.

Side by side

DeepSeek-V3.2-Exp and GPT-5.1 specifications
DeepSeek-V3.2-ExpGPT-5.1
ProviderDeepSeekOpenAI
Noometry Index44.349.0
Released2025-09-292025-11-13
WeightsOpenProprietary
Context window164K400K
Max output66K128K
Input $ / M tokens$0.26$1.25
Output $ / M tokens$0.38$10
Results tracked4963

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

DeepSeek-V3.2-Exp: 46.5 (#65), GPT-5.1: 46.4 (#66)

Coding benchmarks
BenchmarkDeepSeek-V3.2-ExpGPT-5.1
SWE-bench Verified (bash only)70%66%
LMArena WebDev13621395
SciCode38.9%43.3%
WeirdML39.5%60.8%
LMArena Coding14541454
SWE-bench Verified—68%
Aider Polyglot74.2%—
SWE-bench Multilingual59%—
GSO—13.7%
LiveBench Coding—72.5%
ALE-Bench—1,192

Agentic & Tool Use Too close to call

DeepSeek-V3.2-Exp: 32.7 (#59), GPT-5.1: 32.7 (#60)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3.2-ExpGPT-5.1
Terminal-Bench39.6%47.6%
Vending-Bench 21,0341,473
APEX-Agents21.3%—
Berkeley Function Calling Leaderboard56.7%—
TheAgentCompany42.9%—
DeepResearch Bench—42.8%
LMArena Search—1199

Reasoning GPT-5.1 leads

DeepSeek-V3.2-Exp: 22.1 (#208), GPT-5.1: 39.8 (#58)

Reasoning benchmarks
BenchmarkDeepSeek-V3.2-ExpGPT-5.1
ARC-AGI-24%17.6%
ARC-AGI-157%72.8%
CritPt2.9%4.9%
Chess Puzzles14%32%
LMArena Hard Prompts14341457
DTBench87.7%90.1%
LMCA29.1%43.9%
Epoch Capabilities Index146.27149.64
SimpleBench—53.2%
Kagi LLM Benchmark52.2%—
NYT Connections (extended)36.7%—
EnigmaEval—11.2%
Thematic Generalization65%—
LiveBench Reasoning—95.8%
Mystery Game Puzzles—19%
LiveBench Data Analysis—72.1%
ForecastBench—58.1
LiveBench—78.8%

Math GPT-5.1 leads

DeepSeek-V3.2-Exp: 41.7 (#87), GPT-5.1: 52.2 (#51)

Math benchmarks
BenchmarkDeepSeek-V3.2-ExpGPT-5.1
OTIS Mock AIME 2024-202587.8%88.6%
LMArena Math14351447
FrontierMath (Feb 2025 set)22.1%31%
FrontierMath Tier 4 (v1)2.1%12.5%
MathArena Final-Answer Competitions57.7%—
ProofBench8%—
Omni-MATH—46.4%
LiveBench Math—94.5%

Knowledge DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 51.7 (#66), GPT-5.1: 50.6 (#71)

Knowledge benchmarks
BenchmarkDeepSeek-V3.2-ExpGPT-5.1
GPQA Diamond83.4%87.6%
Vectara Hallucination Rate5.3%10.9%
LMArena Expert14361470
Humanity's Last Exam—23.7%
SimpleQA Verified—48%
MMLU-Pro—57.9%
GPQA (HELM)—44.2%

Multimodal Not comparable

DeepSeek-V3.2-Exp: —, GPT-5.1: 44.8 (#19)

Multimodal benchmarks
BenchmarkDeepSeek-V3.2-ExpGPT-5.1
LMArena Vision—1250
VPCT—58.7%
LMArena Document—1403

Multilingual GPT-5.1 leads

DeepSeek-V3.2-Exp: 52.2 (#90), GPT-5.1: 53.8 (#56)

Multilingual benchmarks
BenchmarkDeepSeek-V3.2-ExpGPT-5.1
LMArena Non-English14091431
LMArena Chinese14611495
LMArena French14331450
LMArena German14401438
LMArena Japanese13741453
LMArena Korean13711401
LMArena Russian14241435
LMArena Spanish14401433

Instruction Following GPT-5.1 leads

DeepSeek-V3.2-Exp: 74.5 (#93), GPT-5.1: 83.9 (#1)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.2-ExpGPT-5.1
LMArena Instruction Following14131443
LiveBench Instruction Following—93.3%
IFEval—93.5%

Long Context Too close to call

DeepSeek-V3.2-Exp: 47.6 (#16), GPT-5.1: 47.6 (#14)

Long Context benchmarks
BenchmarkDeepSeek-V3.2-ExpGPT-5.1
CL-bench13.2%23.7%
CL-bench Life9.5%17.3%
LMArena Longer Query14281447
Fiction.LiveBench83.3%—

Writing & Preference GPT-5.1 leads

DeepSeek-V3.2-Exp: 62.4 (#77), GPT-5.1: 64.5 (#55)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.2-ExpGPT-5.1
LMArena Text14251443
LMArena Creative Writing14031427
LMArena Multi-Turn14271450
EQ-Bench Creative Writing1515—
WildBench—86.3%
LiveBench Language—80.2%

Frequently asked questions

Is DeepSeek-V3.2-Exp better than GPT-5.1?

GPT-5.1 is the stronger model overall, scoring 49.0 to 44.3 on the Noometry Index. DeepSeek-V3.2-Exp costs 12× less per token, which makes it the better buy when GPT-5.1's lead doesn't matter for your workload.

Which is cheaper, DeepSeek-V3.2-Exp or GPT-5.1?

DeepSeek-V3.2-Exp is cheaper. It lists at $0.26 per million input tokens and $0.38 per million output tokens; GPT-5.1 lists at $1.25 and $10.

Is DeepSeek-V3.2-Exp or GPT-5.1 better for coding?

They score almost the same on coding (46.5 vs 46.4); test both on your own repository before choosing.

Which has the bigger context window?

GPT-5.1 does, with 400K tokens against 164K.

How many benchmarks do DeepSeek-V3.2-Exp and GPT-5.1 share?

37 benchmarks have published results for both models. DeepSeek-V3.2-Exp has 49 scored results on Noometry and GPT-5.1 has 63.

Related comparisons

Go deeper