Model comparison

DeepSeek-V3 vs DeepSeek-V3.2-Exp

DeepSeek-V3.2-Exp is the stronger model overall, scoring 44.3 to 39.5 on the Noometry Index.

Last verified . 31 shared benchmarks.

DeepSeek-V3 DeepSeek

39.5

Rank #166 Confirmed

DeepSeek-V3.2-Exp DeepSeek

44.3

Rank #78 Confirmed

Summary

  • They share 31 benchmarks with published results for both. DeepSeek-V3 scores higher in 0 categories and DeepSeek-V3.2-Exp in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where DeepSeek-V3.2-Exp leads 51.7 to 37.5.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 37.8% for DeepSeek-V3 and 87.8% for DeepSeek-V3.2-Exp.
  • DeepSeek-V3.2-Exp is cheaper at $0.26 / $0.38 per million input/output tokens, against $0.24 / $0.90 for DeepSeek-V3.

Side by side

DeepSeek-V3 and DeepSeek-V3.2-Exp specifications
DeepSeek-V3DeepSeek-V3.2-Exp
ProviderDeepSeekDeepSeek
Noometry Index39.544.3
Released2024-12-262025-09-29
WeightsOpenOpen
Context window164K164K
Max output164K66K
Input $ / M tokens$0.24$0.26
Output $ / M tokens$0.90$0.38
Results tracked6049

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3.2-Exp leads

DeepSeek-V3: 42.3 (#106), DeepSeek-V3.2-Exp: 46.5 (#65)

Coding benchmarks
BenchmarkDeepSeek-V3DeepSeek-V3.2-Exp
Aider Polyglot55.1%74.2%
SciCode35.8%38.9%
WeirdML36.1%39.5%
LMArena Coding13681454
SWE-bench Verified (bash only)—70%
LMArena WebDev—1362
SWE-bench Multilingual—59%
BigCodeBench Instruct50%—
LiveBench Coding70.9%—
BigCodeBench Complete62.2%—
HumanEval+86.6%—
MBPP+73%—

Agentic & Tool Use Not comparable

DeepSeek-V3: —, DeepSeek-V3.2-Exp: 32.7 (#59)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3DeepSeek-V3.2-Exp
Terminal-Bench—39.6%
APEX-Agents—21.3%
Berkeley Function Calling Leaderboard—56.7%
TheAgentCompany—42.9%
METR Time Horizons49.6%—
Vending-Bench 2—1,034

Reasoning DeepSeek-V3.2-Exp leads

DeepSeek-V3: 20.5 (#236), DeepSeek-V3.2-Exp: 22.1 (#208)

Reasoning benchmarks
BenchmarkDeepSeek-V3DeepSeek-V3.2-Exp
Kagi LLM Benchmark52.3%52.2%
CritPt0%2.9%
LMArena Hard Prompts13651434
DTBench64.8%87.7%
LMCA15.5%29.1%
Epoch Capabilities Index135.94146.27
ARC-AGI-2—4%
SimpleBench27.2%—
NYT Connections (extended)—36.7%
ARC-AGI-1—57%
Chess Puzzles—14%
Thematic Generalization—65%
LiveBench Reasoning65.8%—
LiveBench Data Analysis60.9%—
BIG-Bench Hard87.5%—
ForecastBench59.1—
HellaSwag88.9%—
LiveBench66.9%—
PIQA84.7%—
WinoGrande85.2%—

Math DeepSeek-V3.2-Exp leads

DeepSeek-V3: 32.1 (#219), DeepSeek-V3.2-Exp: 41.7 (#87)

Math benchmarks
BenchmarkDeepSeek-V3DeepSeek-V3.2-Exp
OTIS Mock AIME 2024-202537.8%87.8%
LMArena Math13731435
FrontierMath (Feb 2025 set)1.7%22.1%
MathArena Final-Answer Competitions—57.7%
ProofBench—8%
Omni-MATH40.3%—
LiveBench Math73.5%—
MATH Level 575.5%—
FrontierMath Tier 4 (v1)—2.1%

Knowledge DeepSeek-V3.2-Exp leads

DeepSeek-V3: 37.5 (#155), DeepSeek-V3.2-Exp: 51.7 (#66)

Knowledge benchmarks
BenchmarkDeepSeek-V3DeepSeek-V3.2-Exp
GPQA Diamond67.6%83.4%
Vectara Hallucination Rate6.1%5.3%
LMArena Expert13511436
MMLU-Pro72.3%—
Confabulations26.1%—
GPQA (HELM)53.8%—
ARC (AI2) Challenge95.3%—
MMLU87.2%—
TriviaQA82.9%—

Multilingual DeepSeek-V3.2-Exp leads

DeepSeek-V3: 48.5 (#143), DeepSeek-V3.2-Exp: 52.2 (#90)

Multilingual benchmarks
BenchmarkDeepSeek-V3DeepSeek-V3.2-Exp
LMArena Non-English13581409
LMArena Chinese13911461
LMArena French13851433
LMArena German13741440
LMArena Japanese13331374
LMArena Korean13191371
LMArena Russian13731424
LMArena Spanish13581440

Instruction Following DeepSeek-V3.2-Exp leads

DeepSeek-V3: 72.8 (#130), DeepSeek-V3.2-Exp: 74.5 (#93)

Instruction Following benchmarks
BenchmarkDeepSeek-V3DeepSeek-V3.2-Exp
LMArena Instruction Following13451413
LiveBench Instruction Following81.5%—
IFEval83.2%—

Long Context DeepSeek-V3.2-Exp leads

DeepSeek-V3: 34.0 (#253), DeepSeek-V3.2-Exp: 47.6 (#16)

Long Context benchmarks
BenchmarkDeepSeek-V3DeepSeek-V3.2-Exp
Fiction.LiveBench50%83.3%
LMArena Longer Query13521428
CL-bench—13.2%
CL-bench Life—9.5%

Writing & Preference DeepSeek-V3.2-Exp leads

DeepSeek-V3: 57.4 (#130), DeepSeek-V3.2-Exp: 62.4 (#77)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3DeepSeek-V3.2-Exp
LMArena Text13751425
LMArena Creative Writing13641403
EQ-Bench Creative Writing14721515
LMArena Multi-Turn13891427
Short-Story Creative Writing77%—
WildBench83%—
LiveBench Language49.1%—

Frequently asked questions

Is DeepSeek-V3 better than DeepSeek-V3.2-Exp?

DeepSeek-V3.2-Exp is the stronger model overall, scoring 44.3 to 39.5 on the Noometry Index.

Which is cheaper, DeepSeek-V3 or DeepSeek-V3.2-Exp?

DeepSeek-V3.2-Exp is cheaper. It lists at $0.26 per million input tokens and $0.38 per million output tokens; DeepSeek-V3 lists at $0.24 and $0.90.

Is DeepSeek-V3 or DeepSeek-V3.2-Exp better for coding?

DeepSeek-V3.2-Exp scores higher on coding benchmarks: 46.5 versus 42.3 in the Noometry coding category.

Which has the bigger context window?

Both accept 164K tokens.

How many benchmarks do DeepSeek-V3 and DeepSeek-V3.2-Exp share?

31 benchmarks have published results for both models. DeepSeek-V3 has 60 scored results on Noometry and DeepSeek-V3.2-Exp has 49.

Related comparisons

Go deeper