Model comparison

DBRX vs Grok-2 (Dec 2024)

Grok-2 (Dec 2024) is the stronger model overall, scoring 33.7 to 29.4 on the Noometry Index.

Last verified . 19 shared benchmarks.

DBRX Databricks

29.4

Rank #311 Confirmed

Grok-2 (Dec 2024) xAI

33.7

Rank #239 Confirmed

Summary

  • They share 19 benchmarks with published results for both. DBRX scores higher in 2 categories and Grok-2 (Dec 2024) in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Grok-2 (Dec 2024) leads 48.6 to 33.6.
  • The biggest single-benchmark swing is MATH Level 5: 11.7% for DBRX and 63.5% for Grok-2 (Dec 2024).
  • DBRX has downloadable open weights; the other is API-only.

Side by side

DBRX and Grok-2 (Dec 2024) specifications
DBRXGrok-2 (Dec 2024)
ProviderDatabricksxAI
Noometry Index29.433.7
Released2024-03-272024-08-13
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2134

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

DBRX: 32.9 (#266), Grok-2 (Dec 2024): 33.3 (#258)

Coding benchmarks
BenchmarkDBRXGrok-2 (Dec 2024)
LMArena Coding11321287
WeirdML—22.2%
LiveBench Coding—46.4%
HumanEval+70.1%—
MBPP+55.8%—

Reasoning DBRX leads

DBRX: 21.4 (#222), Grok-2 (Dec 2024): 16.9 (#299)

Reasoning benchmarks
BenchmarkDBRXGrok-2 (Dec 2024)
LMArena Hard Prompts11131272
SimpleBench—22.7%
LiveBench Reasoning—54.8%
DTBench—65.2%
LiveBench Data Analysis—54.5%
Epoch Capabilities Index—130.48
LiveBench—54.3%

Math DBRX leads

DBRX: 24.3 (#269), Grok-2 (Dec 2024): 20.8 (#284)

Math benchmarks
BenchmarkDBRXGrok-2 (Dec 2024)
LMArena Math11451283
MATH Level 511.7%63.5%
OTIS Mock AIME 2024-2025—11.5%
LiveBench Math—54.9%
FrontierMath (Feb 2025 set)—0.7%

Knowledge Grok-2 (Dec 2024) leads

DBRX: 14.9 (#294), Grok-2 (Dec 2024): 29.8 (#233)

Knowledge benchmarks
BenchmarkDBRXGrok-2 (Dec 2024)
GPQA Diamond32.9%53.8%
LMArena Expert10761254
Confabulations—20.1%

Multilingual Grok-2 (Dec 2024) leads

DBRX: 29.3 (#268), Grok-2 (Dec 2024): 43.1 (#188)

Multilingual benchmarks
BenchmarkDBRXGrok-2 (Dec 2024)
LMArena Non-English10711282
LMArena Chinese10681289
LMArena French10961318
LMArena German10571287
LMArena Japanese9901244
LMArena Korean9931237
LMArena Russian10781286
LMArena Spanish10641281

Instruction Following Grok-2 (Dec 2024) leads

DBRX: 57.5 (#270), Grok-2 (Dec 2024): 66.9 (#202)

Instruction Following benchmarks
BenchmarkDBRXGrok-2 (Dec 2024)
LMArena Instruction Following11121270
LiveBench Instruction Following—69.6%

Long Context Grok-2 (Dec 2024) leads

DBRX: 33.7 (#258), Grok-2 (Dec 2024): 38.8 (#190)

Long Context benchmarks
BenchmarkDBRXGrok-2 (Dec 2024)
LMArena Longer Query11121276

Writing & Preference Grok-2 (Dec 2024) leads

DBRX: 33.6 (#275), Grok-2 (Dec 2024): 48.6 (#198)

Writing & Preference benchmarks
BenchmarkDBRXGrok-2 (Dec 2024)
LMArena Text11191305
LMArena Creative Writing11041284
LMArena Multi-Turn11111290
Short-Story Creative Writing—63.6%
LiveBench Language—45.6%

Frequently asked questions

Is DBRX better than Grok-2 (Dec 2024)?

Grok-2 (Dec 2024) is the stronger model overall, scoring 33.7 to 29.4 on the Noometry Index.

Is DBRX or Grok-2 (Dec 2024) better for coding?

They score almost the same on coding (32.9 vs 33.3); test both on your own repository before choosing.

How many benchmarks do DBRX and Grok-2 (Dec 2024) share?

19 benchmarks have published results for both models. DBRX has 21 scored results on Noometry and Grok-2 (Dec 2024) has 34.

Related comparisons

Go deeper