Model comparison

DBRX vs Llama2 70b Steerlm Chat

Llama2 70b Steerlm Chat is the stronger model overall, scoring 31.8 to 29.4 on the Noometry Index.

Last verified . 9 shared benchmarks.

DBRX Databricks

29.4

Rank #311 Confirmed

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Summary

  • They share 9 benchmarks with published results for both. DBRX scores higher in 6 categories and Llama2 70b Steerlm Chat in 1 category; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where Llama2 70b Steerlm Chat leads 31.3 to 24.3.

Side by side

DBRX and Llama2 70b Steerlm Chat specifications
DBRXLlama2 70b Steerlm Chat
ProviderDatabricksNVIDIA
Noometry Index29.431.8
Released2024-03-27—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked219

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DBRX leads

DBRX: 32.9 (#266), Llama2 70b Steerlm Chat: 29.9 (#300)

Coding benchmarks
BenchmarkDBRXLlama2 70b Steerlm Chat
LMArena Coding11321025
HumanEval+70.1%—
MBPP+55.8%—

Reasoning DBRX leads

DBRX: 21.4 (#222), Llama2 70b Steerlm Chat: 20.0 (#246)

Reasoning benchmarks
BenchmarkDBRXLlama2 70b Steerlm Chat
LMArena Hard Prompts11131047

Math Llama2 70b Steerlm Chat leads

DBRX: 24.3 (#269), Llama2 70b Steerlm Chat: 31.3 (#226)

Math benchmarks
BenchmarkDBRXLlama2 70b Steerlm Chat
LMArena Math11451072
MATH Level 511.7%—

Knowledge Not comparable

DBRX: 14.9 (#294), Llama2 70b Steerlm Chat: —

Knowledge benchmarks
BenchmarkDBRXLlama2 70b Steerlm Chat
GPQA Diamond32.9%—
LMArena Expert1076—

Multilingual Too close to call

DBRX: 29.3 (#268), Llama2 70b Steerlm Chat: 28.8 (#270)

Multilingual benchmarks
BenchmarkDBRXLlama2 70b Steerlm Chat
LMArena Non-English10711063
LMArena Chinese1068—
LMArena French1096—
LMArena German1057—
LMArena Japanese990—
LMArena Korean993—
LMArena Russian1078—
LMArena Spanish1064—

Instruction Following DBRX leads

DBRX: 57.5 (#270), Llama2 70b Steerlm Chat: 54.2 (#279)

Instruction Following benchmarks
BenchmarkDBRXLlama2 70b Steerlm Chat
LMArena Instruction Following11121060

Long Context DBRX leads

DBRX: 33.7 (#258), Llama2 70b Steerlm Chat: 30.4 (#288)

Long Context benchmarks
BenchmarkDBRXLlama2 70b Steerlm Chat
LMArena Longer Query1112998

Writing & Preference DBRX leads

DBRX: 33.6 (#275), Llama2 70b Steerlm Chat: 31.6 (#283)

Writing & Preference benchmarks
BenchmarkDBRXLlama2 70b Steerlm Chat
LMArena Text11191098
LMArena Creative Writing11041091
LMArena Multi-Turn11111058

Frequently asked questions

Is DBRX better than Llama2 70b Steerlm Chat?

Llama2 70b Steerlm Chat is the stronger model overall, scoring 31.8 to 29.4 on the Noometry Index.

Is DBRX or Llama2 70b Steerlm Chat better for coding?

DBRX scores higher on coding benchmarks: 32.9 versus 29.9 in the Noometry coding category.

How many benchmarks do DBRX and Llama2 70b Steerlm Chat share?

9 benchmarks have published results for both models. DBRX has 21 scored results on Noometry and Llama2 70b Steerlm Chat has 9.

Related comparisons

Go deeper