Model comparison

DeepSeek-V3.2-Exp vs Gemini 2.0 Flash (Feb 2025)

DeepSeek-V3.2-Exp is the stronger model overall, scoring 44.3 to 35.1 on the Noometry Index.

Last verified . 30 shared benchmarks.

DeepSeek-V3.2-Exp DeepSeek

44.3

Rank #78 Confirmed

Gemini 2.0 Flash (Feb 2025) Google

35.1

Rank #228 Confirmed

Summary

  • They share 30 benchmarks with published results for both. DeepSeek-V3.2-Exp scores higher in 9 categories and Gemini 2.0 Flash (Feb 2025) in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where DeepSeek-V3.2-Exp leads 51.7 to 32.0.
  • The biggest single-benchmark swing is SWE-bench Verified (bash only): 70% for DeepSeek-V3.2-Exp and 13.5% for Gemini 2.0 Flash (Feb 2025).
  • DeepSeek-V3.2-Exp has downloadable open weights; the other is API-only.

Side by side

DeepSeek-V3.2-Exp and Gemini 2.0 Flash (Feb 2025) specifications
DeepSeek-V3.2-ExpGemini 2.0 Flash (Feb 2025)
ProviderDeepSeekGoogle
Noometry Index44.335.1
Released2025-09-292024-12-06
WeightsOpenProprietary
Context window164K—
Max output66K—
Input $ / M tokens$0.26—
Output $ / M tokens$0.38—
Results tracked4954

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 46.5 (#65), Gemini 2.0 Flash (Feb 2025): 28.4 (#315)

Coding benchmarks
BenchmarkDeepSeek-V3.2-ExpGemini 2.0 Flash (Feb 2025)
SWE-bench Verified (bash only)70%13.5%
Aider Polyglot74.2%38.2%
WeirdML39.5%25.8%
LMArena Coding14541350
LMArena WebDev1362—
SWE-bench Multilingual59%—
SciCode38.9%—
BigCodeBench Instruct—45.9%
LiveBench Coding—63.4%
BigCodeBench Complete—59.9%
CadEval—30%

Agentic & Tool Use DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 32.7 (#59), Gemini 2.0 Flash (Feb 2025): 28.1 (#92)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3.2-ExpGemini 2.0 Flash (Feb 2025)
TheAgentCompany42.9%11.4%
Terminal-Bench39.6%—
APEX-Agents21.3%—
Berkeley Function Calling Leaderboard56.7%—
Vending-Bench 21,034—

Reasoning DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 22.1 (#208), Gemini 2.0 Flash (Feb 2025): 15.2 (#318)

Reasoning benchmarks
BenchmarkDeepSeek-V3.2-ExpGemini 2.0 Flash (Feb 2025)
ARC-AGI-24%1.3%
Kagi LLM Benchmark52.2%37.8%
LMArena Hard Prompts14341346
DTBench87.7%63.2%
Epoch Capabilities Index146.27135.36
SimpleBench—31.1%
NYT Connections (extended)36.7%—
ARC-AGI-157%—
CritPt2.9%—
Chess Puzzles14%—
EnigmaEval—1.1%
Thematic Generalization65%—
LiveBench Reasoning—78.2%
LiveBench Data Analysis—69.4%
LMCA29.1%—
LiveBench—66.9%

Math DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 41.7 (#87), Gemini 2.0 Flash (Feb 2025): 37.9 (#146)

Math benchmarks
BenchmarkDeepSeek-V3.2-ExpGemini 2.0 Flash (Feb 2025)
OTIS Mock AIME 2024-202587.8%57.8%
LMArena Math14351352
FrontierMath (Feb 2025 set)22.1%1.7%
MathArena Final-Answer Competitions57.7%—
ProofBench8%—
Omni-MATH—45.9%
LiveBench Math—75.8%
MATH Level 5—82.2%
FrontierMath Tier 4 (v1)2.1%—

Knowledge DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 51.7 (#66), Gemini 2.0 Flash (Feb 2025): 32.0 (#213)

Knowledge benchmarks
BenchmarkDeepSeek-V3.2-ExpGemini 2.0 Flash (Feb 2025)
GPQA Diamond83.4%64.1%
LMArena Expert14361339
Humanity's Last Exam—6.6%
MMLU-Pro—73.7%
Confabulations—12.4%
Vectara Hallucination Rate5.3%—
GPQA (HELM)—55.6%
MMLU—79.7%

Multimodal Not comparable

DeepSeek-V3.2-Exp: —, Gemini 2.0 Flash (Feb 2025): 36.5 (#79)

Multimodal benchmarks
BenchmarkDeepSeek-V3.2-ExpGemini 2.0 Flash (Feb 2025)
LMArena Vision—1158
GeoBench—77%

Multilingual DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 52.2 (#90), Gemini 2.0 Flash (Feb 2025): 47.4 (#149)

Multilingual benchmarks
BenchmarkDeepSeek-V3.2-ExpGemini 2.0 Flash (Feb 2025)
LMArena Non-English14091342
LMArena Chinese14611373
LMArena French14331391
LMArena German14401353
LMArena Japanese13741294
LMArena Korean13711313
LMArena Russian14241351
LMArena Spanish14401363

Instruction Following Too close to call

DeepSeek-V3.2-Exp: 74.5 (#93), Gemini 2.0 Flash (Feb 2025): 74.4 (#97)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.2-ExpGemini 2.0 Flash (Feb 2025)
LMArena Instruction Following14131336
LiveBench Instruction Following—85.8%
IFEval—84.1%

Long Context DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 47.6 (#16), Gemini 2.0 Flash (Feb 2025): 38.1 (#203)

Long Context benchmarks
BenchmarkDeepSeek-V3.2-ExpGemini 2.0 Flash (Feb 2025)
Fiction.LiveBench83.3%61.1%
LMArena Longer Query14281344
CL-bench13.2%—
CL-bench Life9.5%—

Writing & Preference DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 62.4 (#77), Gemini 2.0 Flash (Feb 2025): 49.5 (#190)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.2-ExpGemini 2.0 Flash (Feb 2025)
LMArena Text14251354
LMArena Creative Writing14031340
EQ-Bench Creative Writing15151128
LMArena Multi-Turn14271350
Short-Story Creative Writing—73.8%
WildBench—80%
LiveBench Language—51.3%

Frequently asked questions

Is DeepSeek-V3.2-Exp better than Gemini 2.0 Flash (Feb 2025)?

DeepSeek-V3.2-Exp is the stronger model overall, scoring 44.3 to 35.1 on the Noometry Index.

Is DeepSeek-V3.2-Exp or Gemini 2.0 Flash (Feb 2025) better for coding?

DeepSeek-V3.2-Exp scores higher on coding benchmarks: 46.5 versus 28.4 in the Noometry coding category.

How many benchmarks do DeepSeek-V3.2-Exp and Gemini 2.0 Flash (Feb 2025) share?

30 benchmarks have published results for both models. DeepSeek-V3.2-Exp has 49 scored results on Noometry and Gemini 2.0 Flash (Feb 2025) has 54.

Related comparisons

Go deeper