Model comparison

DeepSeek-V3 vs Grok-3 mini

Grok-3 mini is the stronger model overall, scoring 41.2 to 39.5 on the Noometry Index.

Last verified . 33 shared benchmarks.

DeepSeek-V3 DeepSeek

39.5

Rank #166 Confirmed

Grok-3 mini xAI

41.2

Rank #141 Confirmed

Summary

  • They share 33 benchmarks with published results for both. DeepSeek-V3 scores higher in 4 categories and Grok-3 mini in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Grok-3 mini leads 42.1 to 32.1.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 37.8% for DeepSeek-V3 and 77.8% for Grok-3 mini.
  • DeepSeek-V3 has downloadable open weights; the other is API-only.

Side by side

DeepSeek-V3 and Grok-3 mini specifications
DeepSeek-V3Grok-3 mini
ProviderDeepSeekxAI
Noometry Index39.541.2
Released2024-12-262025-04-09
WeightsOpenProprietary
Context window164K—
Max output164K—
Input $ / M tokens$0.24—
Output $ / M tokens$0.90—
Results tracked6035

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3 leads

DeepSeek-V3: 42.3 (#106), Grok-3 mini: 40.8 (#131)

Coding benchmarks
BenchmarkDeepSeek-V3Grok-3 mini
Aider Polyglot55.1%49.3%
WeirdML36.1%42.6%
LMArena Coding13681379
SciCode35.8%—
BigCodeBench Instruct50%—
LiveBench Coding70.9%—
BigCodeBench Complete62.2%—
HumanEval+86.6%—
MBPP+73%—

Agentic & Tool Use Not comparable

DeepSeek-V3: —, Grok-3 mini: —

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3Grok-3 mini
METR Time Horizons49.6%—

Reasoning DeepSeek-V3 leads

DeepSeek-V3: 20.5 (#236), Grok-3 mini: 13.6 (#334)

Reasoning benchmarks
BenchmarkDeepSeek-V3Grok-3 mini
Kagi LLM Benchmark52.3%61.3%
LMArena Hard Prompts13651375
Epoch Capabilities Index135.94140.35
ARC-AGI-2—0.4%
SimpleBench27.2%—
ARC-AGI-1—16.5%
CritPt0%—
LiveBench Reasoning65.8%—
DTBench64.8%—
LiveBench Data Analysis60.9%—
LMCA15.5%—
BIG-Bench Hard87.5%—
ForecastBench59.1—
HellaSwag88.9%—
LiveBench66.9%—
PIQA84.7%—
WinoGrande85.2%—

Math Grok-3 mini leads

DeepSeek-V3: 32.1 (#219), Grok-3 mini: 42.1 (#85)

Math benchmarks
BenchmarkDeepSeek-V3Grok-3 mini
OTIS Mock AIME 2024-202537.8%77.8%
Omni-MATH40.3%31.8%
LMArena Math13731386
MATH Level 575.5%90.9%
FrontierMath (Feb 2025 set)1.7%5.9%
LiveBench Math73.5%—

Knowledge Grok-3 mini leads

DeepSeek-V3: 37.5 (#155), Grok-3 mini: 46.4 (#81)

Knowledge benchmarks
BenchmarkDeepSeek-V3Grok-3 mini
GPQA Diamond67.6%76.3%
MMLU-Pro72.3%79.9%
Confabulations26.1%10.8%
GPQA (HELM)53.8%67.5%
LMArena Expert13511395
Vectara Hallucination Rate6.1%—
ARC (AI2) Challenge95.3%—
MMLU87.2%—
TriviaQA82.9%—

Multilingual Too close to call

DeepSeek-V3: 48.5 (#143), Grok-3 mini: 48.1 (#145)

Multilingual benchmarks
BenchmarkDeepSeek-V3Grok-3 mini
LMArena Non-English13581352
LMArena Chinese13911387
LMArena French13851357
LMArena German13741349
LMArena Japanese13331342
LMArena Korean13191335
LMArena Russian13731353
LMArena Spanish13581381

Instruction Following Grok-3 mini leads

DeepSeek-V3: 72.8 (#130), Grok-3 mini: 78.5 (#9)

Instruction Following benchmarks
BenchmarkDeepSeek-V3Grok-3 mini
IFEval83.2%95.1%
LMArena Instruction Following13451357
LiveBench Instruction Following81.5%—

Long Context Grok-3 mini leads

DeepSeek-V3: 34.0 (#253), Grok-3 mini: 41.0 (#147)

Long Context benchmarks
BenchmarkDeepSeek-V3Grok-3 mini
Fiction.LiveBench50%66.7%
LMArena Longer Query13521372

Writing & Preference DeepSeek-V3 leads

DeepSeek-V3: 57.4 (#130), Grok-3 mini: 52.5 (#169)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3Grok-3 mini
LMArena Text13751370
LMArena Creative Writing13641342
Short-Story Creative Writing77%73.5%
WildBench83%65.1%
LMArena Multi-Turn13891355
EQ-Bench Creative Writing1472—
LiveBench Language49.1%—

Frequently asked questions

Is DeepSeek-V3 better than Grok-3 mini?

Grok-3 mini is the stronger model overall, scoring 41.2 to 39.5 on the Noometry Index.

Is DeepSeek-V3 or Grok-3 mini better for coding?

DeepSeek-V3 scores higher on coding benchmarks: 42.3 versus 40.8 in the Noometry coding category.

How many benchmarks do DeepSeek-V3 and Grok-3 mini share?

33 benchmarks have published results for both models. DeepSeek-V3 has 60 scored results on Noometry and Grok-3 mini has 35.

Related comparisons

Go deeper