Model comparison

DeepSeek-V3.2-Speciale vs Grok 3

DeepSeek-V3.2-Speciale and Grok 3 score almost the same on the Noometry Index (39.7 vs 39.9), so choose on price, context window or the category you care about most.

Last verified . 3 shared benchmarks.

DeepSeek-V3.2-Speciale DeepSeek

39.7

Rank #162 Reported

Grok 3 xAI

39.9

Rank #157 Confirmed

Summary

  • They share 3 benchmarks with published results for both. DeepSeek-V3.2-Speciale scores higher in 1 category and Grok 3 in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where DeepSeek-V3.2-Speciale leads 32.9 to 13.7.
  • The biggest single-benchmark swing is SimpleBench: 52.6% for DeepSeek-V3.2-Speciale and 36.1% for Grok 3.
  • DeepSeek-V3.2-Speciale has downloadable open weights; the other is API-only.

Side by side

DeepSeek-V3.2-Speciale and Grok 3 specifications
DeepSeek-V3.2-SpecialeGrok 3
ProviderDeepSeekxAI
Noometry Index39.739.9
Released2025-12-012025-04-09
WeightsOpenProprietary
Context window128K—
Max output128K—
Input $ / M tokens$0.58—
Output $ / M tokens$1.68—
Results tracked340

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 3 leads

DeepSeek-V3.2-Speciale: 40.4 (#140), Grok 3: 41.9 (#115)

Coding benchmarks
BenchmarkDeepSeek-V3.2-SpecialeGrok 3
WeirdML46.7%37.2%
Aider Polyglot—53.3%
LMArena Coding—1432

Agentic & Tool Use Not comparable

DeepSeek-V3.2-Speciale: —, Grok 3: 30.5 (#76)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3.2-SpecialeGrok 3
BALROG—29.5%

Reasoning DeepSeek-V3.2-Speciale leads

DeepSeek-V3.2-Speciale: 32.9 (#73), Grok 3: 13.7 (#333)

Reasoning benchmarks
BenchmarkDeepSeek-V3.2-SpecialeGrok 3
SimpleBench52.6%36.1%
ARC-AGI-2—0%
Kagi LLM Benchmark—61.3%
ARC-AGI-1—5.5%
LMArena Hard Prompts—1434
Epoch Capabilities Index—138.33

Math Not comparable

DeepSeek-V3.2-Speciale: —, Grok 3: 38.0 (#145)

Math benchmarks
BenchmarkDeepSeek-V3.2-SpecialeGrok 3
OTIS Mock AIME 2024-2025—55.6%
Omni-MATH—46.4%
LMArena Math—1391
MATH Level 5—88.7%
FrontierMath (Feb 2025 set)—3.8%
FrontierMath Tier 4 (v1)—0%

Knowledge Not comparable

DeepSeek-V3.2-Speciale: —, Grok 3: 46.2 (#82)

Knowledge benchmarks
BenchmarkDeepSeek-V3.2-SpecialeGrok 3
GPQA Diamond—75.8%
MMLU-Pro—78.8%
Confabulations—14.2%
Vectara Hallucination Rate—5.8%
GPQA (HELM)—65%
LMArena Expert—1421

Multilingual Not comparable

DeepSeek-V3.2-Speciale: —, Grok 3: 52.3 (#87)

Multilingual benchmarks
BenchmarkDeepSeek-V3.2-SpecialeGrok 3
LMArena Non-English—1410
LMArena Chinese—1448
LMArena French—1460
LMArena German—1431
LMArena Japanese—1387
LMArena Korean—1373
LMArena Russian—1416
LMArena Spanish—1417

Instruction Following Not comparable

DeepSeek-V3.2-Speciale: —, Grok 3: 75.0 (#73)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.2-SpecialeGrok 3
IFEval—88.4%
LMArena Instruction Following—1409

Long Context Not comparable

DeepSeek-V3.2-Speciale: —, Grok 3: 38.7 (#192)

Long Context benchmarks
BenchmarkDeepSeek-V3.2-SpecialeGrok 3
Fiction.LiveBench—58.3%
LMArena Longer Query—1439

Writing & Preference Grok 3 leads

DeepSeek-V3.2-Speciale: 46.0 (#222), Grok 3: 55.8 (#141)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.2-SpecialeGrok 3
EQ-Bench Creative Writing12761186
LMArena Text—1426
LMArena Creative Writing—1414
Short-Story Creative Writing—76.4%
WildBench—84.9%
LMArena Multi-Turn—1425

Frequently asked questions

Is DeepSeek-V3.2-Speciale better than Grok 3?

DeepSeek-V3.2-Speciale and Grok 3 score almost the same on the Noometry Index (39.7 vs 39.9), so choose on price, context window or the category you care about most.

Is DeepSeek-V3.2-Speciale or Grok 3 better for coding?

Grok 3 scores higher on coding benchmarks: 41.9 versus 40.4 in the Noometry coding category.

How many benchmarks do DeepSeek-V3.2-Speciale and Grok 3 share?

3 benchmarks have published results for both models. DeepSeek-V3.2-Speciale has 3 scored results on Noometry and Grok 3 has 40.

Related comparisons

Go deeper