Model comparison

DeepSeek-V3.2-Speciale vs Gemini 2.5 Flash

DeepSeek-V3.2-Speciale and Gemini 2.5 Flash score almost the same on the Noometry Index (39.7 vs 39.3), so choose on price, context window or the category you care about most.

Last verified . 3 shared benchmarks.

DeepSeek-V3.2-Speciale DeepSeek

39.7

Rank #162 Reported

Gemini 2.5 Flash Google

39.3

Rank #170 Confirmed

Summary

  • They share 3 benchmarks with published results for both. DeepSeek-V3.2-Speciale scores higher in 2 categories and Gemini 2.5 Flash in 1 category; 3 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where DeepSeek-V3.2-Speciale leads 32.9 to 18.1.
  • The biggest single-benchmark swing is SimpleBench: 52.6% for DeepSeek-V3.2-Speciale and 41.2% for Gemini 2.5 Flash.
  • Both cost about the same: $0.58 input and $1.68 output per million tokens.
  • Gemini 2.5 Flash accepts more context: 1.05M tokens versus 128K.
  • DeepSeek-V3.2-Speciale has downloadable open weights; the other is API-only.

Side by side

DeepSeek-V3.2-Speciale and Gemini 2.5 Flash specifications
DeepSeek-V3.2-SpecialeGemini 2.5 Flash
ProviderDeepSeekGoogle
Noometry Index39.739.3
Released2025-12-012025-04-17
WeightsOpenProprietary
Context window128K1.05M
Max output128K66K
Input $ / M tokens$0.58$0.30
Output $ / M tokens$1.68$2.50
Results tracked354

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3.2-Speciale leads

DeepSeek-V3.2-Speciale: 40.4 (#140), Gemini 2.5 Flash: 35.8 (#220)

Coding benchmarks
BenchmarkDeepSeek-V3.2-SpecialeGemini 2.5 Flash
WeirdML46.7%41.9%
SWE-bench Verified (bash only)—28.7%
Aider Polyglot—55.1%
LMArena Coding—1424
ALE-Bench—661.88

Agentic & Tool Use Not comparable

DeepSeek-V3.2-Speciale: —, Gemini 2.5 Flash: 30.8 (#74)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3.2-SpecialeGemini 2.5 Flash
Terminal-Bench—17.1%
Berkeley Function Calling Leaderboard—56.2%
TheAgentCompany—41.1%
BALROG—33.5%
Vending-Bench 2—548.84

Reasoning DeepSeek-V3.2-Speciale leads

DeepSeek-V3.2-Speciale: 32.9 (#73), Gemini 2.5 Flash: 18.1 (#286)

Reasoning benchmarks
BenchmarkDeepSeek-V3.2-SpecialeGemini 2.5 Flash
SimpleBench52.6%41.2%
ARC-AGI-2—2.5%
Kagi LLM Benchmark—56.8%
ARC-AGI-1—33.3%
CritPt—1.1%
EnigmaEval—2.7%
LMArena Hard Prompts—1422
DTBench—76.5%
LMCA—27.5%
Epoch Capabilities Index—143.03
ForecastBench—60.6

Math Not comparable

DeepSeek-V3.2-Speciale: —, Gemini 2.5 Flash: 39.9 (#98)

Math benchmarks
BenchmarkDeepSeek-V3.2-SpecialeGemini 2.5 Flash
OTIS Mock AIME 2024-2025—73.1%
Omni-MATH—38.5%
LMArena Math—1415
FrontierMath (Feb 2025 set)—4.8%
FrontierMath Tier 4 (v1)—4.2%

Knowledge Not comparable

DeepSeek-V3.2-Speciale: —, Gemini 2.5 Flash: 36.4 (#168)

Knowledge benchmarks
BenchmarkDeepSeek-V3.2-SpecialeGemini 2.5 Flash
Humanity's Last Exam—12.1%
MMLU-Pro—63.9%
Confabulations—16.8%
Vectara Hallucination Rate—7.8%
GPQA (HELM)—39%
LMArena Expert—1426

Multimodal Not comparable

DeepSeek-V3.2-Speciale: —, Gemini 2.5 Flash: 41.8 (#32)

Multimodal benchmarks
BenchmarkDeepSeek-V3.2-SpecialeGemini 2.5 Flash
LMArena Vision—1253
GeoBench—76%
VPCT—46.2%
SpatialViz-Bench—36.9%

Multilingual Not comparable

DeepSeek-V3.2-Speciale: —, Gemini 2.5 Flash: 52.3 (#88)

Multilingual benchmarks
BenchmarkDeepSeek-V3.2-SpecialeGemini 2.5 Flash
LMArena Non-English—1409
LMArena Chinese—1450
LMArena French—1433
LMArena German—1418
LMArena Japanese—1405
LMArena Korean—1385
LMArena Russian—1415
LMArena Spanish—1421

Instruction Following Not comparable

DeepSeek-V3.2-Speciale: —, Gemini 2.5 Flash: 75.7 (#54)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.2-SpecialeGemini 2.5 Flash
IFEval—89.8%
LMArena Instruction Following—1405

Long Context Not comparable

DeepSeek-V3.2-Speciale: —, Gemini 2.5 Flash: 47.5 (#17)

Long Context benchmarks
BenchmarkDeepSeek-V3.2-SpecialeGemini 2.5 Flash
Fiction.LiveBench—77.8%
LMArena Longer Query—1419

Writing & Preference Gemini 2.5 Flash leads

DeepSeek-V3.2-Speciale: 46.0 (#222), Gemini 2.5 Flash: 53.8 (#157)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.2-SpecialeGemini 2.5 Flash
EQ-Bench Creative Writing12761137
LMArena Text—1417
LMArena Creative Writing—1400
Short-Story Creative Writing—76.5%
WildBench—81.7%
LMArena Multi-Turn—1408

Frequently asked questions

Is DeepSeek-V3.2-Speciale better than Gemini 2.5 Flash?

DeepSeek-V3.2-Speciale and Gemini 2.5 Flash score almost the same on the Noometry Index (39.7 vs 39.3), so choose on price, context window or the category you care about most.

Which is cheaper, DeepSeek-V3.2-Speciale or Gemini 2.5 Flash?

Gemini 2.5 Flash is cheaper. It lists at $0.30 per million input tokens and $2.50 per million output tokens; DeepSeek-V3.2-Speciale lists at $0.58 and $1.68.

Is DeepSeek-V3.2-Speciale or Gemini 2.5 Flash better for coding?

DeepSeek-V3.2-Speciale scores higher on coding benchmarks: 40.4 versus 35.8 in the Noometry coding category.

Which has the bigger context window?

Gemini 2.5 Flash does, with 1.05M tokens against 128K.

How many benchmarks do DeepSeek-V3.2-Speciale and Gemini 2.5 Flash share?

3 benchmarks have published results for both models. DeepSeek-V3.2-Speciale has 3 scored results on Noometry and Gemini 2.5 Flash has 54.

Related comparisons

Go deeper