Model comparison

DBRX vs Qwen-14B

Qwen-14B is the stronger model overall, scoring 31.4 to 29.4 on the Noometry Index.

Last verified . 10 shared benchmarks.

DBRX Databricks

29.4

Rank #311 Confirmed

Qwen-14B Alibaba (Qwen)

31.4

Rank #275 Confirmed

Summary

  • They share 10 benchmarks with published results for both. DBRX scores higher in 6 categories and Qwen-14B in 1 category; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen-14B leads 31.2 to 24.3.

Side by side

DBRX and Qwen-14B specifications
DBRXQwen-14B
ProviderDatabricksAlibaba (Qwen)
Noometry Index29.431.4
Released2024-03-272023-09-24
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2118

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DBRX leads

DBRX: 32.9 (#266), Qwen-14B: 31.2 (#288)

Coding benchmarks
BenchmarkDBRXQwen-14B
LMArena Coding11321071
HumanEval+70.1%—
MBPP+55.8%—

Reasoning DBRX leads

DBRX: 21.4 (#222), Qwen-14B: 19.6 (#257)

Reasoning benchmarks
BenchmarkDBRXQwen-14B
LMArena Hard Prompts11131027
BIG-Bench Hard—55%
Epoch Capabilities Index—113.03
LAMBADA—71.1%
PIQA—79.9%

Math Qwen-14B leads

DBRX: 24.3 (#269), Qwen-14B: 31.2 (#227)

Math benchmarks
BenchmarkDBRXQwen-14B
LMArena Math11451068
MATH Level 511.7%—
GSM8K—61.3%

Knowledge Not comparable

DBRX: 14.9 (#294), Qwen-14B: —

Knowledge benchmarks
BenchmarkDBRXQwen-14B
GPQA Diamond32.9%—
LMArena Expert1076—
ARC (AI2) Challenge—84.4%
BoolQ—86.2%
MMLU—66.3%

Multilingual DBRX leads

DBRX: 29.3 (#268), Qwen-14B: 27.5 (#275)

Multilingual benchmarks
BenchmarkDBRXQwen-14B
LMArena Non-English10711041
LMArena Chinese10681077
LMArena French1096—
LMArena German1057—
LMArena Japanese990—
LMArena Korean993—
LMArena Russian1078—
LMArena Spanish1064—

Instruction Following DBRX leads

DBRX: 57.5 (#270), Qwen-14B: 52.4 (#289)

Instruction Following benchmarks
BenchmarkDBRXQwen-14B
LMArena Instruction Following11121031

Long Context DBRX leads

DBRX: 33.7 (#258), Qwen-14B: 31.3 (#280)

Long Context benchmarks
BenchmarkDBRXQwen-14B
LMArena Longer Query11121028

Writing & Preference DBRX leads

DBRX: 33.6 (#275), Qwen-14B: 27.6 (#299)

Writing & Preference benchmarks
BenchmarkDBRXQwen-14B
LMArena Text11191051
LMArena Creative Writing11041028
LMArena Multi-Turn11111022

Frequently asked questions

Is DBRX better than Qwen-14B?

Qwen-14B is the stronger model overall, scoring 31.4 to 29.4 on the Noometry Index.

Is DBRX or Qwen-14B better for coding?

DBRX scores higher on coding benchmarks: 32.9 versus 31.2 in the Noometry coding category.

How many benchmarks do DBRX and Qwen-14B share?

10 benchmarks have published results for both models. DBRX has 21 scored results on Noometry and Qwen-14B has 18.

Related comparisons

Go deeper