Model comparison

DeepSeek-V3.2-Exp vs Grok 3

DeepSeek-V3.2-Exp is the stronger model overall, scoring 44.3 to 39.9 on the Noometry Index.

Last verified . 30 shared benchmarks.

DeepSeek-V3.2-Exp DeepSeek

44.3

Rank #78 Confirmed

Grok 3 xAI

39.9

Rank #157 Confirmed

Summary

  • They share 30 benchmarks with published results for both. DeepSeek-V3.2-Exp scores higher in 7 categories and Grok 3 in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in long context, where DeepSeek-V3.2-Exp leads 47.6 to 38.7.
  • The biggest single-benchmark swing is ARC-AGI-1: 57% for DeepSeek-V3.2-Exp and 5.5% for Grok 3.
  • DeepSeek-V3.2-Exp has downloadable open weights; the other is API-only.

Side by side

DeepSeek-V3.2-Exp and Grok 3 specifications
DeepSeek-V3.2-ExpGrok 3
ProviderDeepSeekxAI
Noometry Index44.339.9
Released2025-09-292025-04-09
WeightsOpenProprietary
Context window164K—
Max output66K—
Input $ / M tokens$0.26—
Output $ / M tokens$0.38—
Results tracked4940

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 46.5 (#65), Grok 3: 41.9 (#115)

Coding benchmarks
BenchmarkDeepSeek-V3.2-ExpGrok 3
Aider Polyglot74.2%53.3%
WeirdML39.5%37.2%
LMArena Coding14541432
SWE-bench Verified (bash only)70%—
LMArena WebDev1362—
SWE-bench Multilingual59%—
SciCode38.9%—

Agentic & Tool Use DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 32.7 (#59), Grok 3: 30.5 (#76)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3.2-ExpGrok 3
Terminal-Bench39.6%—
APEX-Agents21.3%—
Berkeley Function Calling Leaderboard56.7%—
TheAgentCompany42.9%—
BALROG—29.5%
Vending-Bench 21,034—

Reasoning DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 22.1 (#208), Grok 3: 13.7 (#333)

Reasoning benchmarks
BenchmarkDeepSeek-V3.2-ExpGrok 3
ARC-AGI-24%0%
Kagi LLM Benchmark52.2%61.3%
ARC-AGI-157%5.5%
LMArena Hard Prompts14341434
Epoch Capabilities Index146.27138.33
SimpleBench—36.1%
NYT Connections (extended)36.7%—
CritPt2.9%—
Chess Puzzles14%—
Thematic Generalization65%—
DTBench87.7%—
LMCA29.1%—

Math DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 41.7 (#87), Grok 3: 38.0 (#145)

Math benchmarks
BenchmarkDeepSeek-V3.2-ExpGrok 3
OTIS Mock AIME 2024-202587.8%55.6%
LMArena Math14351391
FrontierMath (Feb 2025 set)22.1%3.8%
FrontierMath Tier 4 (v1)2.1%0%
MathArena Final-Answer Competitions57.7%—
ProofBench8%—
Omni-MATH—46.4%
MATH Level 5—88.7%

Knowledge DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 51.7 (#66), Grok 3: 46.2 (#82)

Knowledge benchmarks
BenchmarkDeepSeek-V3.2-ExpGrok 3
GPQA Diamond83.4%75.8%
Vectara Hallucination Rate5.3%5.8%
LMArena Expert14361421
MMLU-Pro—78.8%
Confabulations—14.2%
GPQA (HELM)—65%

Multilingual Too close to call

DeepSeek-V3.2-Exp: 52.2 (#90), Grok 3: 52.3 (#87)

Multilingual benchmarks
BenchmarkDeepSeek-V3.2-ExpGrok 3
LMArena Non-English14091410
LMArena Chinese14611448
LMArena French14331460
LMArena German14401431
LMArena Japanese13741387
LMArena Korean13711373
LMArena Russian14241416
LMArena Spanish14401417

Instruction Following Too close to call

DeepSeek-V3.2-Exp: 74.5 (#93), Grok 3: 75.0 (#73)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.2-ExpGrok 3
LMArena Instruction Following14131409
IFEval—88.4%

Long Context DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 47.6 (#16), Grok 3: 38.7 (#192)

Long Context benchmarks
BenchmarkDeepSeek-V3.2-ExpGrok 3
Fiction.LiveBench83.3%58.3%
LMArena Longer Query14281439
CL-bench13.2%—
CL-bench Life9.5%—

Writing & Preference DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 62.4 (#77), Grok 3: 55.8 (#141)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.2-ExpGrok 3
LMArena Text14251426
LMArena Creative Writing14031414
EQ-Bench Creative Writing15151186
LMArena Multi-Turn14271425
Short-Story Creative Writing—76.4%
WildBench—84.9%

Frequently asked questions

Is DeepSeek-V3.2-Exp better than Grok 3?

DeepSeek-V3.2-Exp is the stronger model overall, scoring 44.3 to 39.9 on the Noometry Index.

Is DeepSeek-V3.2-Exp or Grok 3 better for coding?

DeepSeek-V3.2-Exp scores higher on coding benchmarks: 46.5 versus 41.9 in the Noometry coding category.

How many benchmarks do DeepSeek-V3.2-Exp and Grok 3 share?

30 benchmarks have published results for both models. DeepSeek-V3.2-Exp has 49 scored results on Noometry and Grok 3 has 40.

Related comparisons

Go deeper