Model comparison
DeepSeek-V3.2-Speciale vs Gemini 2.5 Flash
DeepSeek-V3.2-Speciale and Gemini 2.5 Flash score almost the same on the Noometry Index (39.7 vs 39.3), so choose on price, context window or the category you care about most.
Last verified . 3 shared benchmarks.
Summary
- They share 3 benchmarks with published results for both. DeepSeek-V3.2-Speciale scores higher in 2 categories and Gemini 2.5 Flash in 1 category; 3 gaps are clear of the uncertainty.
- The widest gap is in reasoning, where DeepSeek-V3.2-Speciale leads 32.9 to 18.1.
- The biggest single-benchmark swing is SimpleBench: 52.6% for DeepSeek-V3.2-Speciale and 41.2% for Gemini 2.5 Flash.
- Both cost about the same: $0.58 input and $1.68 output per million tokens.
- Gemini 2.5 Flash accepts more context: 1.05M tokens versus 128K.
- DeepSeek-V3.2-Speciale has downloadable open weights; the other is API-only.
Side by side
| DeepSeek-V3.2-Speciale | Gemini 2.5 Flash | |
|---|---|---|
| Provider | DeepSeek | |
| Noometry Index | 39.7 | 39.3 |
| Released | 2025-12-01 | 2025-04-17 |
| Weights | Open | Proprietary |
| Context window | 128K | 1.05M |
| Max output | 128K | 66K |
| Input $ / M tokens | $0.58 | $0.30 |
| Output $ / M tokens | $1.68 | $2.50 |
| Results tracked | 3 | 54 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding DeepSeek-V3.2-Speciale leads
DeepSeek-V3.2-Speciale: 40.4 (#140), Gemini 2.5 Flash: 35.8 (#220)
| Benchmark | DeepSeek-V3.2-Speciale | Gemini 2.5 Flash |
|---|---|---|
| WeirdML | 46.7% | 41.9% |
| SWE-bench Verified (bash only) | — | 28.7% |
| Aider Polyglot | — | 55.1% |
| LMArena Coding | — | 1424 |
| ALE-Bench | — | 661.88 |
Agentic & Tool Use Not comparable
DeepSeek-V3.2-Speciale: —, Gemini 2.5 Flash: 30.8 (#74)
| Benchmark | DeepSeek-V3.2-Speciale | Gemini 2.5 Flash |
|---|---|---|
| Terminal-Bench | — | 17.1% |
| Berkeley Function Calling Leaderboard | — | 56.2% |
| TheAgentCompany | — | 41.1% |
| BALROG | — | 33.5% |
| Vending-Bench 2 | — | 548.84 |
Reasoning DeepSeek-V3.2-Speciale leads
DeepSeek-V3.2-Speciale: 32.9 (#73), Gemini 2.5 Flash: 18.1 (#286)
| Benchmark | DeepSeek-V3.2-Speciale | Gemini 2.5 Flash |
|---|---|---|
| SimpleBench | 52.6% | 41.2% |
| ARC-AGI-2 | — | 2.5% |
| Kagi LLM Benchmark | — | 56.8% |
| ARC-AGI-1 | — | 33.3% |
| CritPt | — | 1.1% |
| EnigmaEval | — | 2.7% |
| LMArena Hard Prompts | — | 1422 |
| DTBench | — | 76.5% |
| LMCA | — | 27.5% |
| Epoch Capabilities Index | — | 143.03 |
| ForecastBench | — | 60.6 |
Math Not comparable
DeepSeek-V3.2-Speciale: —, Gemini 2.5 Flash: 39.9 (#98)
| Benchmark | DeepSeek-V3.2-Speciale | Gemini 2.5 Flash |
|---|---|---|
| OTIS Mock AIME 2024-2025 | — | 73.1% |
| Omni-MATH | — | 38.5% |
| LMArena Math | — | 1415 |
| FrontierMath (Feb 2025 set) | — | 4.8% |
| FrontierMath Tier 4 (v1) | — | 4.2% |
Knowledge Not comparable
DeepSeek-V3.2-Speciale: —, Gemini 2.5 Flash: 36.4 (#168)
| Benchmark | DeepSeek-V3.2-Speciale | Gemini 2.5 Flash |
|---|---|---|
| Humanity's Last Exam | — | 12.1% |
| MMLU-Pro | — | 63.9% |
| Confabulations | — | 16.8% |
| Vectara Hallucination Rate | — | 7.8% |
| GPQA (HELM) | — | 39% |
| LMArena Expert | — | 1426 |
Multimodal Not comparable
DeepSeek-V3.2-Speciale: —, Gemini 2.5 Flash: 41.8 (#32)
| Benchmark | DeepSeek-V3.2-Speciale | Gemini 2.5 Flash |
|---|---|---|
| LMArena Vision | — | 1253 |
| GeoBench | — | 76% |
| VPCT | — | 46.2% |
| SpatialViz-Bench | — | 36.9% |
Multilingual Not comparable
DeepSeek-V3.2-Speciale: —, Gemini 2.5 Flash: 52.3 (#88)
| Benchmark | DeepSeek-V3.2-Speciale | Gemini 2.5 Flash |
|---|---|---|
| LMArena Non-English | — | 1409 |
| LMArena Chinese | — | 1450 |
| LMArena French | — | 1433 |
| LMArena German | — | 1418 |
| LMArena Japanese | — | 1405 |
| LMArena Korean | — | 1385 |
| LMArena Russian | — | 1415 |
| LMArena Spanish | — | 1421 |
Instruction Following Not comparable
DeepSeek-V3.2-Speciale: —, Gemini 2.5 Flash: 75.7 (#54)
| Benchmark | DeepSeek-V3.2-Speciale | Gemini 2.5 Flash |
|---|---|---|
| IFEval | — | 89.8% |
| LMArena Instruction Following | — | 1405 |
Long Context Not comparable
DeepSeek-V3.2-Speciale: —, Gemini 2.5 Flash: 47.5 (#17)
| Benchmark | DeepSeek-V3.2-Speciale | Gemini 2.5 Flash |
|---|---|---|
| Fiction.LiveBench | — | 77.8% |
| LMArena Longer Query | — | 1419 |
Writing & Preference Gemini 2.5 Flash leads
DeepSeek-V3.2-Speciale: 46.0 (#222), Gemini 2.5 Flash: 53.8 (#157)
| Benchmark | DeepSeek-V3.2-Speciale | Gemini 2.5 Flash |
|---|---|---|
| EQ-Bench Creative Writing | 1276 | 1137 |
| LMArena Text | — | 1417 |
| LMArena Creative Writing | — | 1400 |
| Short-Story Creative Writing | — | 76.5% |
| WildBench | — | 81.7% |
| LMArena Multi-Turn | — | 1408 |
Frequently asked questions
Is DeepSeek-V3.2-Speciale better than Gemini 2.5 Flash?
DeepSeek-V3.2-Speciale and Gemini 2.5 Flash score almost the same on the Noometry Index (39.7 vs 39.3), so choose on price, context window or the category you care about most.
Which is cheaper, DeepSeek-V3.2-Speciale or Gemini 2.5 Flash?
Gemini 2.5 Flash is cheaper. It lists at $0.30 per million input tokens and $2.50 per million output tokens; DeepSeek-V3.2-Speciale lists at $0.58 and $1.68.
Is DeepSeek-V3.2-Speciale or Gemini 2.5 Flash better for coding?
DeepSeek-V3.2-Speciale scores higher on coding benchmarks: 40.4 versus 35.8 in the Noometry coding category.
Which has the bigger context window?
Gemini 2.5 Flash does, with 1.05M tokens against 128K.
How many benchmarks do DeepSeek-V3.2-Speciale and Gemini 2.5 Flash share?
3 benchmarks have published results for both models. DeepSeek-V3.2-Speciale has 3 scored results on Noometry and Gemini 2.5 Flash has 54.