Model comparison

DeepSeek-V3.1-Terminus vs Grok 4.1 Fast

DeepSeek-V3.1-Terminus is the stronger model overall, scoring 43.1 to 41.4 on the Noometry Index. Grok 4.1 Fast costs 1.6× less per token, which makes it the better buy when DeepSeek-V3.1-Terminus's lead doesn't matter for your workload.

Last verified . 12 shared benchmarks.

DeepSeek-V3.1-Terminus DeepSeek

43.1

Rank #97 Confirmed

Grok 4.1 Fast xAI

41.4

Rank #136 Confirmed

Summary

  • They share 12 benchmarks with published results for both. DeepSeek-V3.1-Terminus scores higher in 6 categories and Grok 4.1 Fast in 1 category; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.1 Fast leads 43.4 to 26.4.
  • The biggest single-benchmark swing is DTBench: 81.3% for DeepSeek-V3.1-Terminus and 87.7% for Grok 4.1 Fast.
  • Grok 4.1 Fast is cheaper at $0.20 / $0.50 per million input/output tokens, against $0.27 / $1 for DeepSeek-V3.1-Terminus.
  • DeepSeek-V3.1-Terminus accepts more context: 164K tokens versus 128K.
  • DeepSeek-V3.1-Terminus has downloadable open weights; the other is API-only.

Side by side

DeepSeek-V3.1-Terminus and Grok 4.1 Fast specifications
DeepSeek-V3.1-TerminusGrok 4.1 Fast
ProviderDeepSeekxAI
Noometry Index43.141.4
Released2025-09-222025-06-27
WeightsOpenProprietary
Context window164K128K
Max output147K30K
Input $ / M tokens$0.27$0.20
Output $ / M tokens$1$0.50
Results tracked1632

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3.1-Terminus leads

DeepSeek-V3.1-Terminus: 42.0 (#113), Grok 4.1 Fast: 34.1 (#245)

Coding benchmarks
BenchmarkDeepSeek-V3.1-TerminusGrok 4.1 Fast
LMArena Coding14261411
ALE-Bench745.17394.93
LMArena WebDev—1242
SciCode40.6%—

Agentic & Tool Use Not comparable

DeepSeek-V3.1-Terminus: —, Grok 4.1 Fast: 36.3 (#39)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3.1-TerminusGrok 4.1 Fast
Berkeley Function Calling Leaderboard—69.6%
τ²-bench Banking—13.1%
LMArena Search—1171
Vending-Bench 2—1,107

Reasoning Grok 4.1 Fast leads

DeepSeek-V3.1-Terminus: 26.4 (#133), Grok 4.1 Fast: 43.4 (#49)

Reasoning benchmarks
BenchmarkDeepSeek-V3.1-TerminusGrok 4.1 Fast
LMArena Hard Prompts14261407
DTBench81.3%87.7%
SimpleBench—56%
Kagi LLM Benchmark57.4%—
NYT Connections (extended)—87.4%
CritPt1.7%—
LMCA28.6%—
ForecastBench—61

Math DeepSeek-V3.1-Terminus leads

DeepSeek-V3.1-Terminus: 38.5 (#137), Grok 4.1 Fast: 31.9 (#221)

Math benchmarks
BenchmarkDeepSeek-V3.1-TerminusGrok 4.1 Fast
LMArena Math14021408
MathArena Final-Answer Competitions—60.9%
ProofBench—4%

Knowledge Not comparable

DeepSeek-V3.1-Terminus: —, Grok 4.1 Fast: 33.1 (#207)

Knowledge benchmarks
BenchmarkDeepSeek-V3.1-TerminusGrok 4.1 Fast
Vectara Hallucination Rate—17.8%
LMArena Expert—1399

Multimodal Not comparable

DeepSeek-V3.1-Terminus: —, Grok 4.1 Fast: 37.0 (#76)

Multimodal benchmarks
BenchmarkDeepSeek-V3.1-TerminusGrok 4.1 Fast
LMArena Vision—1201

Multilingual DeepSeek-V3.1-Terminus leads

DeepSeek-V3.1-Terminus: 52.1 (#92), Grok 4.1 Fast: 51.0 (#114)

Multilingual benchmarks
BenchmarkDeepSeek-V3.1-TerminusGrok 4.1 Fast
LMArena Non-English14071391
LMArena Russian14361387
LMArena Chinese—1441
LMArena French—1415
LMArena German—1404
LMArena Japanese—1349
LMArena Korean—1361
LMArena Spanish—1413

Instruction Following DeepSeek-V3.1-Terminus leads

DeepSeek-V3.1-Terminus: 74.0 (#106), Grok 4.1 Fast: 72.7 (#133)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.1-TerminusGrok 4.1 Fast
LMArena Instruction Following14041376

Long Context DeepSeek-V3.1-Terminus leads

DeepSeek-V3.1-Terminus: 43.4 (#97), Grok 4.1 Fast: 42.4 (#126)

Long Context benchmarks
BenchmarkDeepSeek-V3.1-TerminusGrok 4.1 Fast
LMArena Longer Query14211390

Writing & Preference DeepSeek-V3.1-Terminus leads

DeepSeek-V3.1-Terminus: 61.0 (#92), Grok 4.1 Fast: 57.2 (#131)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.1-TerminusGrok 4.1 Fast
LMArena Text14191408
LMArena Creative Writing14031394
LMArena Multi-Turn14111389
EQ-Bench Creative Writing—1327

Frequently asked questions

Is DeepSeek-V3.1-Terminus better than Grok 4.1 Fast?

DeepSeek-V3.1-Terminus is the stronger model overall, scoring 43.1 to 41.4 on the Noometry Index. Grok 4.1 Fast costs 1.6× less per token, which makes it the better buy when DeepSeek-V3.1-Terminus's lead doesn't matter for your workload.

Which is cheaper, DeepSeek-V3.1-Terminus or Grok 4.1 Fast?

Grok 4.1 Fast is cheaper. It lists at $0.20 per million input tokens and $0.50 per million output tokens; DeepSeek-V3.1-Terminus lists at $0.27 and $1.

Is DeepSeek-V3.1-Terminus or Grok 4.1 Fast better for coding?

DeepSeek-V3.1-Terminus scores higher on coding benchmarks: 42.0 versus 34.1 in the Noometry coding category.

Which has the bigger context window?

DeepSeek-V3.1-Terminus does, with 164K tokens against 128K.

How many benchmarks do DeepSeek-V3.1-Terminus and Grok 4.1 Fast share?

12 benchmarks have published results for both models. DeepSeek-V3.1-Terminus has 16 scored results on Noometry and Grok 4.1 Fast has 32.

Related comparisons

Go deeper