Model comparison

DBRX vs Grok 4.3

Grok 4.3 is the stronger model overall, scoring 43.8 to 29.4 on the Noometry Index.

Last verified . 18 shared benchmarks.

DBRX Databricks

29.4

Rank #311 Confirmed

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Summary

  • They share 18 benchmarks with published results for both. DBRX scores higher in 0 categories and Grok 4.3 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok 4.3 leads 52.5 to 14.9.
  • The biggest single-benchmark swing is GPQA Diamond: 32.9% for DBRX and 88.8% for Grok 4.3.
  • DBRX has downloadable open weights; the other is API-only.

Side by side

DBRX and Grok 4.3 specifications
DBRXGrok 4.3
ProviderDatabricksxAI
Noometry Index29.443.8
Released2024-03-272026-04-17
WeightsOpenProprietary
Context window—1M
Max output—30K
Input $ / M tokens—$1.25
Output $ / M tokens—$2.50
Results tracked2140

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.3 leads

DBRX: 32.9 (#266), Grok 4.3: 41.6 (#121)

Coding benchmarks
BenchmarkDBRXGrok 4.3
LMArena Coding11321415
LMArena WebDev—1357
SciCode—47.3%
WeirdML—49.9%
ALE-Bench—944.17
HumanEval+70.1%—
MBPP+55.8%—

Agentic & Tool Use Not comparable

DBRX: —, Grok 4.3: 27.7 (#99)

Agentic & Tool Use benchmarks
BenchmarkDBRXGrok 4.3
GDP.pdf—8%
LMArena Search—1165
Vending-Bench 2—35.26

Reasoning Grok 4.3 leads

DBRX: 21.4 (#222), Grok 4.3: 35.9 (#68)

Reasoning benchmarks
BenchmarkDBRXGrok 4.3
LMArena Hard Prompts11131396
NYT Connections (extended)—55.2%
CritPt—8%
Chess Puzzles—25%
DTBench—90.7%
LMCA—38.3%
Epoch Capabilities Index—149.16
ForecastBench—60.3

Math Grok 4.3 leads

DBRX: 24.3 (#269), Grok 4.3: 46.0 (#74)

Math benchmarks
BenchmarkDBRXGrok 4.3
LMArena Math11451388
FrontierMath (Tiers 1-3)—42.8%
FrontierMath Tier 4—14.6%
OTIS Mock AIME 2024-2025—93.3%
ProofBench—11%
MATH Level 511.7%—

Knowledge Grok 4.3 leads

DBRX: 14.9 (#294), Grok 4.3: 52.5 (#62)

Knowledge benchmarks
BenchmarkDBRXGrok 4.3
GPQA Diamond32.9%88.8%
LMArena Expert10761385
SimpleQA Verified—33.2%

Multimodal Not comparable

DBRX: —, Grok 4.3: 31.6 (#104)

Multimodal benchmarks
BenchmarkDBRXGrok 4.3
LMArena Vision—1229
Blueprint-Bench 2—0%

Multilingual Grok 4.3 leads

DBRX: 29.3 (#268), Grok 4.3: 50.5 (#120)

Multilingual benchmarks
BenchmarkDBRXGrok 4.3
LMArena Non-English10711385
LMArena Chinese10681422
LMArena French10961412
LMArena German10571395
LMArena Japanese9901379
LMArena Korean9931356
LMArena Russian10781399
LMArena Spanish10641398

Instruction Following Grok 4.3 leads

DBRX: 57.5 (#270), Grok 4.3: 72.1 (#140)

Instruction Following benchmarks
BenchmarkDBRXGrok 4.3
LMArena Instruction Following11121366

Long Context Grok 4.3 leads

DBRX: 33.7 (#258), Grok 4.3: 42.5 (#123)

Long Context benchmarks
BenchmarkDBRXGrok 4.3
LMArena Longer Query11121393

Writing & Preference Grok 4.3 leads

DBRX: 33.6 (#275), Grok 4.3: 58.5 (#118)

Writing & Preference benchmarks
BenchmarkDBRXGrok 4.3
LMArena Text11191397
LMArena Creative Writing11041380
LMArena Multi-Turn11111406
EQ-Bench 4—1075

Frequently asked questions

Is DBRX better than Grok 4.3?

Grok 4.3 is the stronger model overall, scoring 43.8 to 29.4 on the Noometry Index.

Is DBRX or Grok 4.3 better for coding?

Grok 4.3 scores higher on coding benchmarks: 41.6 versus 32.9 in the Noometry coding category.

How many benchmarks do DBRX and Grok 4.3 share?

18 benchmarks have published results for both models. DBRX has 21 scored results on Noometry and Grok 4.3 has 40.

Related comparisons

Go deeper