Model comparison

DeepSeek-V3.1 vs Grok 4.3

DeepSeek-V3.1 and Grok 4.3 score almost the same on the Noometry Index (42.8 vs 43.8), so choose on price, context window or the category you care about most.

Last verified . 22 shared benchmarks.

DeepSeek-V3.1 DeepSeek

42.8

Rank #108 Confirmed

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Summary

  • They share 22 benchmarks with published results for both. DeepSeek-V3.1 scores higher in 3 categories and Grok 4.3 in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok 4.3 leads 52.5 to 43.7.
  • The biggest single-benchmark swing is LMCA: 24.3% for DeepSeek-V3.1 and 38.3% for Grok 4.3.
  • DeepSeek-V3.1 is cheaper at $0.25 / $0.95 per million input/output tokens, against $1.25 / $2.50 for Grok 4.3.
  • Grok 4.3 accepts more context: 1M tokens versus 164K.
  • DeepSeek-V3.1 has downloadable open weights; the other is API-only.

Side by side

DeepSeek-V3.1 and Grok 4.3 specifications
DeepSeek-V3.1Grok 4.3
ProviderDeepSeekxAI
Noometry Index42.843.8
Released2025-08-212026-04-17
WeightsOpenProprietary
Context window164K1M
Max output8K30K
Input $ / M tokens$0.25$1.25
Output $ / M tokens$0.95$2.50
Results tracked2740

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.3 leads

DeepSeek-V3.1: 40.3 (#144), Grok 4.3: 41.6 (#121)

Coding benchmarks
BenchmarkDeepSeek-V3.1Grok 4.3
WeirdML38.4%49.9%
LMArena Coding14171415
LMArena WebDev—1357
SciCode—47.3%
ALE-Bench—944.17

Agentic & Tool Use Not comparable

DeepSeek-V3.1: —, Grok 4.3: 27.7 (#99)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3.1Grok 4.3
GDP.pdf—8%
LMArena Search—1165
Vending-Bench 2—35.26

Reasoning Grok 4.3 leads

DeepSeek-V3.1: 27.9 (#110), Grok 4.3: 35.9 (#68)

Reasoning benchmarks
BenchmarkDeepSeek-V3.1Grok 4.3
LMArena Hard Prompts14171396
DTBench82.7%90.7%
LMCA24.3%38.3%
Epoch Capabilities Index139.92149.16
ForecastBench5860.3
SimpleBench40%—
Kagi LLM Benchmark53.2%—
NYT Connections (extended)—55.2%
CritPt—8%
Chess Puzzles—25%

Math Grok 4.3 leads

DeepSeek-V3.1: 38.9 (#122), Grok 4.3: 46.0 (#74)

Math benchmarks
BenchmarkDeepSeek-V3.1Grok 4.3
LMArena Math14201388
FrontierMath (Tiers 1-3)—42.8%
FrontierMath Tier 4—14.6%
OTIS Mock AIME 2024-2025—93.3%
ProofBench—11%

Knowledge Grok 4.3 leads

DeepSeek-V3.1: 43.7 (#90), Grok 4.3: 52.5 (#62)

Knowledge benchmarks
BenchmarkDeepSeek-V3.1Grok 4.3
LMArena Expert14051385
GPQA Diamond—88.8%
SimpleQA Verified—33.2%
Vectara Hallucination Rate5.5%—

Multimodal Not comparable

DeepSeek-V3.1: —, Grok 4.3: 31.6 (#104)

Multimodal benchmarks
BenchmarkDeepSeek-V3.1Grok 4.3
LMArena Vision—1229
Blueprint-Bench 2—0%

Multilingual DeepSeek-V3.1 leads

DeepSeek-V3.1: 51.6 (#106), Grok 4.3: 50.5 (#120)

Multilingual benchmarks
BenchmarkDeepSeek-V3.1Grok 4.3
LMArena Non-English14001385
LMArena Chinese14691422
LMArena French14471412
LMArena German14111395
LMArena Japanese13781379
LMArena Korean13371356
LMArena Russian14051399
LMArena Spanish14311398

Instruction Following DeepSeek-V3.1 leads

DeepSeek-V3.1: 73.9 (#110), Grok 4.3: 72.1 (#140)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.1Grok 4.3
LMArena Instruction Following14001366

Long Context Grok 4.3 leads

DeepSeek-V3.1: 36.3 (#232), Grok 4.3: 42.5 (#123)

Long Context benchmarks
BenchmarkDeepSeek-V3.1Grok 4.3
LMArena Longer Query14221393
Fiction.LiveBench52.8%—

Writing & Preference DeepSeek-V3.1 leads

DeepSeek-V3.1: 60.3 (#98), Grok 4.3: 58.5 (#118)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.1Grok 4.3
LMArena Text14201397
LMArena Creative Writing14011380
LMArena Multi-Turn14081406
EQ-Bench Creative Writing1436—
EQ-Bench 4—1075

Frequently asked questions

Is DeepSeek-V3.1 better than Grok 4.3?

DeepSeek-V3.1 and Grok 4.3 score almost the same on the Noometry Index (42.8 vs 43.8), so choose on price, context window or the category you care about most.

Which is cheaper, DeepSeek-V3.1 or Grok 4.3?

DeepSeek-V3.1 is cheaper. It lists at $0.25 per million input tokens and $0.95 per million output tokens; Grok 4.3 lists at $1.25 and $2.50.

Is DeepSeek-V3.1 or Grok 4.3 better for coding?

Grok 4.3 scores higher on coding benchmarks: 41.6 versus 40.3 in the Noometry coding category.

Which has the bigger context window?

Grok 4.3 does, with 1M tokens against 164K.

How many benchmarks do DeepSeek-V3.1 and Grok 4.3 share?

22 benchmarks have published results for both models. DeepSeek-V3.1 has 27 scored results on Noometry and Grok 4.3 has 40.

Related comparisons

Go deeper