Model comparison

DBRX vs Llama 13b

DBRX is the stronger model overall, scoring 29.4 to 24.4 on the Noometry Index.

Last verified . 8 shared benchmarks.

DBRX Databricks

29.4

Rank #311 Confirmed

Llama 13b Meta

24.4

Rank #348 Confirmed

Summary

  • They share 8 benchmarks with published results for both. DBRX scores higher in 5 categories and Llama 13b in 1 category; 6 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where DBRX leads 57.5 to 36.7.

Side by side

DBRX and Llama 13b specifications
DBRXLlama 13b
ProviderDatabricksMeta
Noometry Index29.424.4
Released2024-03-272023-02-24
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2121

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DBRX leads

DBRX: 32.9 (#266), Llama 13b: 21.4 (#337)

Coding benchmarks
BenchmarkDBRXLlama 13b
LMArena Coding1132683
HumanEval+70.1%—
MBPP+55.8%—

Reasoning DBRX leads

DBRX: 21.4 (#222), Llama 13b: 14.0 (#329)

Reasoning benchmarks
BenchmarkDBRXLlama 13b
LMArena Hard Prompts1113728
BIG-Bench Hard—37.9%
Epoch Capabilities Index—100.58
HellaSwag—79.2%
LAMBADA—75.2%
PIQA—80.1%
WinoGrande—73%

Math Llama 13b leads

DBRX: 24.3 (#269), Llama 13b: 26.7 (#256)

Math benchmarks
BenchmarkDBRXLlama 13b
LMArena Math1145838
MATH Level 511.7%—
GSM8K—20.6%

Knowledge Not comparable

DBRX: 14.9 (#294), Llama 13b: —

Knowledge benchmarks
BenchmarkDBRXLlama 13b
GPQA Diamond32.9%—
LMArena Expert1076—
ARC (AI2) Challenge—52.7%
BoolQ—78.7%
MMLU—47.7%
OpenBookQA—56.4%
TriviaQA—77.9%

Multimodal Not comparable

DBRX: —, Llama 13b: —

Multimodal benchmarks
BenchmarkDBRXLlama 13b
ScienceQA—43.3%

Multilingual DBRX leads

DBRX: 29.3 (#268), Llama 13b: 16.6 (#297)

Multilingual benchmarks
BenchmarkDBRXLlama 13b
LMArena Non-English1071819
LMArena Chinese1068—
LMArena French1096—
LMArena German1057—
LMArena Japanese990—
LMArena Korean993—
LMArena Russian1078—
LMArena Spanish1064—

Instruction Following DBRX leads

DBRX: 57.5 (#270), Llama 13b: 36.7 (#305)

Instruction Following benchmarks
BenchmarkDBRXLlama 13b
LMArena Instruction Following1112781

Long Context Not comparable

DBRX: 33.7 (#258), Llama 13b: —

Long Context benchmarks
BenchmarkDBRXLlama 13b
LMArena Longer Query1112—

Writing & Preference DBRX leads

DBRX: 33.6 (#275), Llama 13b: 13.8 (#312)

Writing & Preference benchmarks
BenchmarkDBRXLlama 13b
LMArena Text1119834
LMArena Creative Writing1104794
LMArena Multi-Turn1111753

Frequently asked questions

Is DBRX better than Llama 13b?

DBRX is the stronger model overall, scoring 29.4 to 24.4 on the Noometry Index.

Is DBRX or Llama 13b better for coding?

DBRX scores higher on coding benchmarks: 32.9 versus 21.4 in the Noometry coding category.

How many benchmarks do DBRX and Llama 13b share?

8 benchmarks have published results for both models. DBRX has 21 scored results on Noometry and Llama 13b has 21.

Related comparisons

Go deeper