Model comparison

DBRX vs Llama 2-70B

DBRX is the stronger model overall, scoring 29.4 to 24.4 on the Noometry Index.

Last verified . 19 shared benchmarks.

DBRX Databricks

29.4

Rank #311 Confirmed

Llama 2-70B Meta

24.4

Rank #349 Confirmed

Summary

  • They share 19 benchmarks with published results for both. DBRX scores higher in 8 categories and Llama 2-70B in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where DBRX leads 24.3 to 8.1.
  • The biggest single-benchmark swing is MATH Level 5: 11.7% for DBRX and 3.3% for Llama 2-70B.

Side by side

DBRX and Llama 2-70B specifications
DBRXLlama 2-70B
ProviderDatabricksMeta
Noometry Index29.424.4
Released2024-03-272023-07-18
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2135

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DBRX leads

DBRX: 32.9 (#266), Llama 2-70B: 31.4 (#286)

Coding benchmarks
BenchmarkDBRXLlama 2-70B
LMArena Coding11321079
HumanEval+70.1%—
MBPP+55.8%—

Reasoning DBRX leads

DBRX: 21.4 (#222), Llama 2-70B: 14.4 (#325)

Reasoning benchmarks
BenchmarkDBRXLlama 2-70B
LMArena Hard Prompts11131073
DTBench—41.6%
BIG-Bench Hard—64.9%
CommonsenseQA 2.0—50%
Epoch Capabilities Index—113.79
ForecastBench—51.4
HellaSwag—85.3%
LAMBADA—78.9%
PIQA—82.8%
WinoGrande—80.2%

Math DBRX leads

DBRX: 24.3 (#269), Llama 2-70B: 8.1 (#326)

Math benchmarks
BenchmarkDBRXLlama 2-70B
LMArena Math11451091
MATH Level 511.7%3.3%
OTIS Mock AIME 2024-2025—0%
GSM8K—69.6%

Knowledge DBRX leads

DBRX: 14.9 (#294), Llama 2-70B: 7.4 (#310)

Knowledge benchmarks
BenchmarkDBRXLlama 2-70B
GPQA Diamond32.9%26.3%
LMArena Expert10761039
ARC (AI2) Challenge—78.3%
BoolQ—88.6%
MMLU—69.9%
OpenBookQA—60.2%
TriviaQA—87.6%

Multilingual DBRX leads

DBRX: 29.3 (#268), Llama 2-70B: 27.7 (#274)

Multilingual benchmarks
BenchmarkDBRXLlama 2-70B
LMArena Non-English10711045
LMArena Chinese1068995
LMArena French10961090
LMArena German10571041
LMArena Japanese990927
LMArena Korean993964
LMArena Russian10781083
LMArena Spanish10641143

Instruction Following DBRX leads

DBRX: 57.5 (#270), Llama 2-70B: 54.9 (#278)

Instruction Following benchmarks
BenchmarkDBRXLlama 2-70B
LMArena Instruction Following11121071

Long Context DBRX leads

DBRX: 33.7 (#258), Llama 2-70B: 32.3 (#270)

Long Context benchmarks
BenchmarkDBRXLlama 2-70B
LMArena Longer Query11121062

Writing & Preference DBRX leads

DBRX: 33.6 (#275), Llama 2-70B: 32.3 (#279)

Writing & Preference benchmarks
BenchmarkDBRXLlama 2-70B
LMArena Text11191115
LMArena Creative Writing11041075
LMArena Multi-Turn11111088

Frequently asked questions

Is DBRX better than Llama 2-70B?

DBRX is the stronger model overall, scoring 29.4 to 24.4 on the Noometry Index.

Is DBRX or Llama 2-70B better for coding?

DBRX scores higher on coding benchmarks: 32.9 versus 31.4 in the Noometry coding category.

How many benchmarks do DBRX and Llama 2-70B share?

19 benchmarks have published results for both models. DBRX has 21 scored results on Noometry and Llama 2-70B has 35.

Related comparisons

Go deeper