Model comparison

DBRX vs Qwen2.5 Plus 1127

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 29.4 on the Noometry Index.

Last verified . 14 shared benchmarks.

DBRX Databricks

29.4

Rank #311 Confirmed

Qwen2.5 Plus 1127 Alibaba (Qwen)

38.8

Rank #181 Confirmed

Summary

  • They share 14 benchmarks with published results for both. DBRX scores higher in 0 categories and Qwen2.5 Plus 1127 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen2.5 Plus 1127 leads 35.5 to 14.9.
  • DBRX has downloadable open weights; the other is API-only.

Side by side

DBRX and Qwen2.5 Plus 1127 specifications
DBRXQwen2.5 Plus 1127
ProviderDatabricksAlibaba (Qwen)
Noometry Index29.438.8
Released2024-03-27—
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2114

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 Plus 1127 leads

DBRX: 32.9 (#266), Qwen2.5 Plus 1127: 38.5 (#175)

Coding benchmarks
BenchmarkDBRXQwen2.5 Plus 1127
LMArena Coding11321314
HumanEval+70.1%—
MBPP+55.8%—

Reasoning Qwen2.5 Plus 1127 leads

DBRX: 21.4 (#222), Qwen2.5 Plus 1127: 25.9 (#141)

Reasoning benchmarks
BenchmarkDBRXQwen2.5 Plus 1127
LMArena Hard Prompts11131299

Math Qwen2.5 Plus 1127 leads

DBRX: 24.3 (#269), Qwen2.5 Plus 1127: 36.1 (#174)

Math benchmarks
BenchmarkDBRXQwen2.5 Plus 1127
LMArena Math11451298
MATH Level 511.7%—

Knowledge Qwen2.5 Plus 1127 leads

DBRX: 14.9 (#294), Qwen2.5 Plus 1127: 35.5 (#183)

Knowledge benchmarks
BenchmarkDBRXQwen2.5 Plus 1127
LMArena Expert10761289
GPQA Diamond32.9%—

Multilingual Qwen2.5 Plus 1127 leads

DBRX: 29.3 (#268), Qwen2.5 Plus 1127: 41.9 (#201)

Multilingual benchmarks
BenchmarkDBRXQwen2.5 Plus 1127
LMArena Non-English10711265
LMArena Chinese10681314
LMArena German10571231
LMArena Japanese9901207
LMArena Russian10781271
LMArena French1096—
LMArena Korean993—
LMArena Spanish1064—

Instruction Following Qwen2.5 Plus 1127 leads

DBRX: 57.5 (#270), Qwen2.5 Plus 1127: 67.2 (#199)

Instruction Following benchmarks
BenchmarkDBRXQwen2.5 Plus 1127
LMArena Instruction Following11121275

Long Context Qwen2.5 Plus 1127 leads

DBRX: 33.7 (#258), Qwen2.5 Plus 1127: 39.2 (#184)

Long Context benchmarks
BenchmarkDBRXQwen2.5 Plus 1127
LMArena Longer Query11121292

Writing & Preference Qwen2.5 Plus 1127 leads

DBRX: 33.6 (#275), Qwen2.5 Plus 1127: 49.4 (#192)

Writing & Preference benchmarks
BenchmarkDBRXQwen2.5 Plus 1127
LMArena Text11191299
LMArena Creative Writing11041262
LMArena Multi-Turn11111299

Frequently asked questions

Is DBRX better than Qwen2.5 Plus 1127?

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 29.4 on the Noometry Index.

Is DBRX or Qwen2.5 Plus 1127 better for coding?

Qwen2.5 Plus 1127 scores higher on coding benchmarks: 38.5 versus 32.9 in the Noometry coding category.

How many benchmarks do DBRX and Qwen2.5 Plus 1127 share?

14 benchmarks have published results for both models. DBRX has 21 scored results on Noometry and Qwen2.5 Plus 1127 has 14.

Related comparisons

Go deeper