Model comparison

DeepSeek-R1 vs Grok 3

DeepSeek-R1 is the stronger model overall, scoring 42.3 to 39.9 on the Noometry Index.

Last verified . 38 shared benchmarks.

DeepSeek-R1 DeepSeek

42.3

Rank #115 Confirmed

Grok 3 xAI

39.9

Rank #157 Confirmed

Summary

  • They share 38 benchmarks with published results for both. DeepSeek-R1 scores higher in 7 categories and Grok 3 in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in long context, where DeepSeek-R1 leads 45.4 to 38.7.
  • The biggest single-benchmark swing is Aider Polyglot: 71.4% for DeepSeek-R1 and 53.3% for Grok 3.

Side by side

DeepSeek-R1 and Grok 3 specifications
DeepSeek-R1Grok 3
ProviderDeepSeekxAI
Noometry Index42.339.9
Released2025-01-202025-04-09
WeightsProprietaryProprietary
Context window164K—
Max output64K—
Input $ / M tokens$0.50—
Output $ / M tokens$2.15—
Results tracked5240

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-R1 leads

DeepSeek-R1: 46.3 (#68), Grok 3: 41.9 (#115)

Coding benchmarks
BenchmarkDeepSeek-R1Grok 3
Aider Polyglot71.4%53.3%
WeirdML41.6%37.2%
LMArena Coding14271432
SciCode35.7%—
LiveBench Coding66.7%—
ALE-Bench804.12—
AlgoTune1.7—

Agentic & Tool Use Too close to call

DeepSeek-R1: 30.7 (#75), Grok 3: 30.5 (#76)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-R1Grok 3
BALROG34.9%29.5%
DeepResearch Bench35.1%—
METR Time Horizons53.8%—

Reasoning DeepSeek-R1 leads

DeepSeek-R1: 18.6 (#278), Grok 3: 13.7 (#333)

Reasoning benchmarks
BenchmarkDeepSeek-R1Grok 3
ARC-AGI-21.3%0%
SimpleBench40.8%36.1%
Kagi LLM Benchmark69.4%61.3%
ARC-AGI-121.2%5.5%
LMArena Hard Prompts14161434
Epoch Capabilities Index141.29138.33
CritPt1.1%—
LiveBench Reasoning83.2%—
LiveBench Data Analysis69.8%—
ForecastBench60—
LiveBench71.6%—

Math DeepSeek-R1 leads

DeepSeek-R1: 43.8 (#79), Grok 3: 38.0 (#145)

Math benchmarks
BenchmarkDeepSeek-R1Grok 3
OTIS Mock AIME 2024-202566.4%55.6%
Omni-MATH42.4%46.4%
LMArena Math14001391
MATH Level 596.6%88.7%
LiveBench Math80.7%—
FrontierMath (Feb 2025 set)—3.8%
FrontierMath Tier 4 (v1)—0%

Knowledge Grok 3 leads

DeepSeek-R1: 44.5 (#87), Grok 3: 46.2 (#82)

Knowledge benchmarks
BenchmarkDeepSeek-R1Grok 3
GPQA Diamond76.3%75.8%
MMLU-Pro79.3%78.8%
Confabulations12.7%14.2%
Vectara Hallucination Rate11.3%5.8%
GPQA (HELM)66.6%65%
LMArena Expert13941421

Multilingual Too close to call

DeepSeek-R1: 52.4 (#85), Grok 3: 52.3 (#87)

Multilingual benchmarks
BenchmarkDeepSeek-R1Grok 3
LMArena Non-English14121410
LMArena Chinese14421448
LMArena French14171460
LMArena German14041431
LMArena Japanese13911387
LMArena Korean13601373
LMArena Russian14231416
LMArena Spanish14111417

Instruction Following Grok 3 leads

DeepSeek-R1: 72.0 (#143), Grok 3: 75.0 (#73)

Instruction Following benchmarks
BenchmarkDeepSeek-R1Grok 3
IFEval78.4%88.4%
LMArena Instruction Following13821409
LiveBench Instruction Following80.5%—

Long Context DeepSeek-R1 leads

DeepSeek-R1: 45.4 (#36), Grok 3: 38.7 (#192)

Long Context benchmarks
BenchmarkDeepSeek-R1Grok 3
Fiction.LiveBench75%58.3%
LMArena Longer Query13911439

Writing & Preference DeepSeek-R1 leads

DeepSeek-R1: 61.4 (#88), Grok 3: 55.8 (#141)

Writing & Preference benchmarks
BenchmarkDeepSeek-R1Grok 3
LMArena Text14281426
LMArena Creative Writing14051414
Short-Story Creative Writing83%76.4%
EQ-Bench Creative Writing15001186
WildBench82.8%84.9%
LMArena Multi-Turn14051425
LiveBench Language48.5%—

Frequently asked questions

Is DeepSeek-R1 better than Grok 3?

DeepSeek-R1 is the stronger model overall, scoring 42.3 to 39.9 on the Noometry Index.

Is DeepSeek-R1 or Grok 3 better for coding?

DeepSeek-R1 scores higher on coding benchmarks: 46.3 versus 41.9 in the Noometry coding category.

How many benchmarks do DeepSeek-R1 and Grok 3 share?

38 benchmarks have published results for both models. DeepSeek-R1 has 52 scored results on Noometry and Grok 3 has 40.

Related comparisons

Go deeper