Model comparison

DBRX vs Grok 4.1

Grok 4.1 is the stronger model overall, scoring 41.5 to 29.4 on the Noometry Index.

Last verified . 17 shared benchmarks.

DBRX Databricks

29.4

Rank #311 Confirmed

Grok 4.1 xAI

41.5

Rank #134 Confirmed

Summary

  • They share 17 benchmarks with published results for both. DBRX scores higher in 0 categories and Grok 4.1 in 8 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Grok 4.1 leads 62.4 to 33.6.
  • DBRX has downloadable open weights; the other is API-only.

Side by side

DBRX and Grok 4.1 specifications
DBRXGrok 4.1
ProviderDatabricksxAI
Noometry Index29.441.5
Released2024-03-272025-11-17
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2119

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

DBRX: 32.9 (#266), Grok 4.1: 33.7 (#253)

Coding benchmarks
BenchmarkDBRXGrok 4.1
LMArena Coding11321445
LMArena WebDev—1214
HumanEval+70.1%—
MBPP+55.8%—

Agentic & Tool Use Not comparable

DBRX: —, Grok 4.1: 34.1 (#49)

Agentic & Tool Use benchmarks
BenchmarkDBRXGrok 4.1
Cybench—39%

Reasoning Grok 4.1 leads

DBRX: 21.4 (#222), Grok 4.1: 29.5 (#91)

Reasoning benchmarks
BenchmarkDBRXGrok 4.1
LMArena Hard Prompts11131435

Math Grok 4.1 leads

DBRX: 24.3 (#269), Grok 4.1: 38.9 (#120)

Math benchmarks
BenchmarkDBRXGrok 4.1
LMArena Math11451422
MATH Level 511.7%—

Knowledge Grok 4.1 leads

DBRX: 14.9 (#294), Grok 4.1: 39.5 (#133)

Knowledge benchmarks
BenchmarkDBRXGrok 4.1
LMArena Expert10761417
GPQA Diamond32.9%—

Multilingual Grok 4.1 leads

DBRX: 29.3 (#268), Grok 4.1: 53.4 (#68)

Multilingual benchmarks
BenchmarkDBRXGrok 4.1
LMArena Non-English10711425
LMArena Chinese10681465
LMArena French10961448
LMArena German10571446
LMArena Japanese9901397
LMArena Korean9931407
LMArena Russian10781434
LMArena Spanish10641438

Instruction Following Grok 4.1 leads

DBRX: 57.5 (#270), Grok 4.1: 73.8 (#111)

Instruction Following benchmarks
BenchmarkDBRXGrok 4.1
LMArena Instruction Following11121400

Long Context Grok 4.1 leads

DBRX: 33.7 (#258), Grok 4.1: 43.2 (#100)

Long Context benchmarks
BenchmarkDBRXGrok 4.1
LMArena Longer Query11121416

Writing & Preference Grok 4.1 leads

DBRX: 33.6 (#275), Grok 4.1: 62.4 (#75)

Writing & Preference benchmarks
BenchmarkDBRXGrok 4.1
LMArena Text11191437
LMArena Creative Writing11041411
LMArena Multi-Turn11111437

Frequently asked questions

Is DBRX better than Grok 4.1?

Grok 4.1 is the stronger model overall, scoring 41.5 to 29.4 on the Noometry Index.

Is DBRX or Grok 4.1 better for coding?

They score almost the same on coding (32.9 vs 33.7); test both on your own repository before choosing.

How many benchmarks do DBRX and Grok 4.1 share?

17 benchmarks have published results for both models. DBRX has 21 scored results on Noometry and Grok 4.1 has 19.

Related comparisons

Go deeper