Model comparison

DBRX vs Llama 3-8B

DBRX is the stronger model overall, scoring 29.4 to 25.5 on the Noometry Index.

Last verified . 21 shared benchmarks.

DBRX Databricks

29.4

Rank #311 Confirmed

Llama 3-8B Meta

25.5

Rank #344 Confirmed

Summary

  • They share 21 benchmarks with published results for both. DBRX scores higher in 4 categories and Llama 3-8B in 4 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where DBRX leads 24.3 to 8.8.
  • The biggest single-benchmark swing is GPQA Diamond: 32.9% for DBRX and 26.1% for Llama 3-8B.

Side by side

DBRX and Llama 3-8B specifications
DBRXLlama 3-8B
ProviderDatabricksMeta
Noometry Index29.425.5
Released2024-03-272024-04-18
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2134

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DBRX leads

DBRX: 32.9 (#266), Llama 3-8B: 31.0 (#289)

Coding benchmarks
BenchmarkDBRXLlama 3-8B
LMArena Coding11321152
HumanEval+70.1%56.7%
MBPP+55.8%54.8%
BigCodeBench Instruct—31.9%
BigCodeBench Complete—36.9%

Reasoning DBRX leads

DBRX: 21.4 (#222), Llama 3-8B: 14.3 (#326)

Reasoning benchmarks
BenchmarkDBRXLlama 3-8B
LMArena Hard Prompts11131133
Chess Puzzles—0%
DTBench—43.9%
Adversarial NLI—57.3%
Epoch Capabilities Index—116.45
ForecastBench—58.6
WinoGrande—75.7%

Math DBRX leads

DBRX: 24.3 (#269), Llama 3-8B: 8.8 (#323)

Math benchmarks
BenchmarkDBRXLlama 3-8B
LMArena Math11451151
MATH Level 511.7%6.1%
OTIS Mock AIME 2024-2025—1.9%

Knowledge DBRX leads

DBRX: 14.9 (#294), Llama 3-8B: 7.8 (#308)

Knowledge benchmarks
BenchmarkDBRXLlama 3-8B
GPQA Diamond32.9%26.1%
LMArena Expert10761113
ARC (AI2) Challenge—82.8%
MMLU—68.8%
OpenBookQA—82.6%
TriviaQA—67.7%

Multilingual Llama 3-8B leads

DBRX: 29.3 (#268), Llama 3-8B: 30.8 (#261)

Multilingual benchmarks
BenchmarkDBRXLlama 3-8B
LMArena Non-English10711098
LMArena Chinese10681076
LMArena French10961159
LMArena German10571104
LMArena Japanese990967
LMArena Korean9931004
LMArena Russian10781109
LMArena Spanish10641173

Instruction Following Too close to call

DBRX: 57.5 (#270), Llama 3-8B: 58.4 (#260)

Instruction Following benchmarks
BenchmarkDBRXLlama 3-8B
LMArena Instruction Following11121127

Long Context Too close to call

DBRX: 33.7 (#258), Llama 3-8B: 34.2 (#251)

Long Context benchmarks
BenchmarkDBRXLlama 3-8B
LMArena Longer Query11121128

Writing & Preference Llama 3-8B leads

DBRX: 33.6 (#275), Llama 3-8B: 37.5 (#256)

Writing & Preference benchmarks
BenchmarkDBRXLlama 3-8B
LMArena Text11191166
LMArena Creative Writing11041150
LMArena Multi-Turn11111152

Frequently asked questions

Is DBRX better than Llama 3-8B?

DBRX is the stronger model overall, scoring 29.4 to 25.5 on the Noometry Index.

Is DBRX or Llama 3-8B better for coding?

DBRX scores higher on coding benchmarks: 32.9 versus 31.0 in the Noometry coding category.

How many benchmarks do DBRX and Llama 3-8B share?

21 benchmarks have published results for both models. DBRX has 21 scored results on Noometry and Llama 3-8B has 34.

Related comparisons

Go deeper