Model comparison

DeepSeek-V3.1-Terminus vs Grok 4.3

DeepSeek-V3.1-Terminus and Grok 4.3 score almost the same on the Noometry Index (43.1 vs 43.8), so choose on price, context window or the category you care about most.

Last verified . 15 shared benchmarks.

DeepSeek-V3.1-Terminus DeepSeek

43.1

Rank #97 Confirmed

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Summary

  • They share 15 benchmarks with published results for both. DeepSeek-V3.1-Terminus scores higher in 5 categories and Grok 4.3 in 2 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.3 leads 35.9 to 26.4.
  • The biggest single-benchmark swing is LMCA: 28.6% for DeepSeek-V3.1-Terminus and 38.3% for Grok 4.3.
  • DeepSeek-V3.1-Terminus is cheaper at $0.27 / $1 per million input/output tokens, against $1.25 / $2.50 for Grok 4.3.
  • Grok 4.3 accepts more context: 1M tokens versus 164K.
  • DeepSeek-V3.1-Terminus has downloadable open weights; the other is API-only.

Side by side

DeepSeek-V3.1-Terminus and Grok 4.3 specifications
DeepSeek-V3.1-TerminusGrok 4.3
ProviderDeepSeekxAI
Noometry Index43.143.8
Released2025-09-222026-04-17
WeightsOpenProprietary
Context window164K1M
Max output147K30K
Input $ / M tokens$0.27$1.25
Output $ / M tokens$1$2.50
Results tracked1640

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

DeepSeek-V3.1-Terminus: 42.0 (#113), Grok 4.3: 41.6 (#121)

Coding benchmarks
BenchmarkDeepSeek-V3.1-TerminusGrok 4.3
SciCode40.6%47.3%
LMArena Coding14261415
ALE-Bench745.17944.17
LMArena WebDev—1357
WeirdML—49.9%

Agentic & Tool Use Not comparable

DeepSeek-V3.1-Terminus: —, Grok 4.3: 27.7 (#99)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3.1-TerminusGrok 4.3
GDP.pdf—8%
LMArena Search—1165
Vending-Bench 2—35.26

Reasoning Grok 4.3 leads

DeepSeek-V3.1-Terminus: 26.4 (#133), Grok 4.3: 35.9 (#68)

Reasoning benchmarks
BenchmarkDeepSeek-V3.1-TerminusGrok 4.3
CritPt1.7%8%
LMArena Hard Prompts14261396
DTBench81.3%90.7%
LMCA28.6%38.3%
Kagi LLM Benchmark57.4%—
NYT Connections (extended)—55.2%
Chess Puzzles—25%
Epoch Capabilities Index—149.16
ForecastBench—60.3

Math Grok 4.3 leads

DeepSeek-V3.1-Terminus: 38.5 (#137), Grok 4.3: 46.0 (#74)

Math benchmarks
BenchmarkDeepSeek-V3.1-TerminusGrok 4.3
LMArena Math14021388
FrontierMath (Tiers 1-3)—42.8%
FrontierMath Tier 4—14.6%
OTIS Mock AIME 2024-2025—93.3%
ProofBench—11%

Knowledge Not comparable

DeepSeek-V3.1-Terminus: —, Grok 4.3: 52.5 (#62)

Knowledge benchmarks
BenchmarkDeepSeek-V3.1-TerminusGrok 4.3
GPQA Diamond—88.8%
SimpleQA Verified—33.2%
LMArena Expert—1385

Multimodal Not comparable

DeepSeek-V3.1-Terminus: —, Grok 4.3: 31.6 (#104)

Multimodal benchmarks
BenchmarkDeepSeek-V3.1-TerminusGrok 4.3
LMArena Vision—1229
Blueprint-Bench 2—0%

Multilingual DeepSeek-V3.1-Terminus leads

DeepSeek-V3.1-Terminus: 52.1 (#92), Grok 4.3: 50.5 (#120)

Multilingual benchmarks
BenchmarkDeepSeek-V3.1-TerminusGrok 4.3
LMArena Non-English14071385
LMArena Russian14361399
LMArena Chinese—1422
LMArena French—1412
LMArena German—1395
LMArena Japanese—1379
LMArena Korean—1356
LMArena Spanish—1398

Instruction Following DeepSeek-V3.1-Terminus leads

DeepSeek-V3.1-Terminus: 74.0 (#106), Grok 4.3: 72.1 (#140)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.1-TerminusGrok 4.3
LMArena Instruction Following14041366

Long Context Too close to call

DeepSeek-V3.1-Terminus: 43.4 (#97), Grok 4.3: 42.5 (#123)

Long Context benchmarks
BenchmarkDeepSeek-V3.1-TerminusGrok 4.3
LMArena Longer Query14211393

Writing & Preference DeepSeek-V3.1-Terminus leads

DeepSeek-V3.1-Terminus: 61.0 (#92), Grok 4.3: 58.5 (#118)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.1-TerminusGrok 4.3
LMArena Text14191397
LMArena Creative Writing14031380
LMArena Multi-Turn14111406
EQ-Bench 4—1075

Frequently asked questions

Is DeepSeek-V3.1-Terminus better than Grok 4.3?

DeepSeek-V3.1-Terminus and Grok 4.3 score almost the same on the Noometry Index (43.1 vs 43.8), so choose on price, context window or the category you care about most.

Which is cheaper, DeepSeek-V3.1-Terminus or Grok 4.3?

DeepSeek-V3.1-Terminus is cheaper. It lists at $0.27 per million input tokens and $1 per million output tokens; Grok 4.3 lists at $1.25 and $2.50.

Is DeepSeek-V3.1-Terminus or Grok 4.3 better for coding?

They score almost the same on coding (42.0 vs 41.6); test both on your own repository before choosing.

Which has the bigger context window?

Grok 4.3 does, with 1M tokens against 164K.

How many benchmarks do DeepSeek-V3.1-Terminus and Grok 4.3 share?

15 benchmarks have published results for both models. DeepSeek-V3.1-Terminus has 16 scored results on Noometry and Grok 4.3 has 40.

Related comparisons

Go deeper