Model comparison

Claude 2.1 vs DBRX

DBRX is the stronger model overall, scoring 29.4 to 25.2 on the Noometry Index.

Last verified . 1 shared benchmarks.

Claude 2.1 Anthropic

25.2

Rank #345 Reported

DBRX Databricks

29.4

Rank #311 Confirmed

Summary

  • They share 1 benchmark with published results for both. Claude 2.1 scores higher in 2 categories and DBRX in 2 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in math, where DBRX leads 24.3 to 10.2.
  • DBRX has downloadable open weights; the other is API-only.

Side by side

Claude 2.1 and DBRX specifications
Claude 2.1DBRX
ProviderAnthropicDatabricks
Noometry Index25.229.4
Released2023-11-212024-03-27
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked721

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DBRX leads

Claude 2.1: 26.2 (#327), DBRX: 32.9 (#266)

Coding benchmarks
BenchmarkClaude 2.1DBRX
WeirdML7.1%—
LMArena Coding—1132
HumanEval+—70.1%
MBPP+—55.8%

Reasoning Too close to call

Claude 2.1: 21.4 (#221), DBRX: 21.4 (#222)

Reasoning benchmarks
BenchmarkClaude 2.1DBRX
LMArena Hard Prompts—1113
DTBench51%—
Epoch Capabilities Index119.27—
ForecastBench54.2—

Math DBRX leads

Claude 2.1: 10.2 (#315), DBRX: 24.3 (#269)

Math benchmarks
BenchmarkClaude 2.1DBRX
OTIS Mock AIME 2024-20251.9%—
LMArena Math—1145
MATH Level 5—11.7%

Knowledge Too close to call

Claude 2.1: 15.4 (#292), DBRX: 14.9 (#294)

Knowledge benchmarks
BenchmarkClaude 2.1DBRX
GPQA Diamond33%32.9%
LMArena Expert—1076
MMLU73.5%—

Multilingual Not comparable

Claude 2.1: —, DBRX: 29.3 (#268)

Multilingual benchmarks
BenchmarkClaude 2.1DBRX
LMArena Non-English—1071
LMArena Chinese—1068
LMArena French—1096
LMArena German—1057
LMArena Japanese—990
LMArena Korean—993
LMArena Russian—1078
LMArena Spanish—1064

Instruction Following Not comparable

Claude 2.1: —, DBRX: 57.5 (#270)

Instruction Following benchmarks
BenchmarkClaude 2.1DBRX
LMArena Instruction Following—1112

Long Context Not comparable

Claude 2.1: —, DBRX: 33.7 (#258)

Long Context benchmarks
BenchmarkClaude 2.1DBRX
LMArena Longer Query—1112

Writing & Preference Not comparable

Claude 2.1: —, DBRX: 33.6 (#275)

Writing & Preference benchmarks
BenchmarkClaude 2.1DBRX
LMArena Text—1119
LMArena Creative Writing—1104
LMArena Multi-Turn—1111

Frequently asked questions

Is Claude 2.1 better than DBRX?

DBRX is the stronger model overall, scoring 29.4 to 25.2 on the Noometry Index.

Is Claude 2.1 or DBRX better for coding?

DBRX scores higher on coding benchmarks: 32.9 versus 26.2 in the Noometry coding category.

How many benchmarks do Claude 2.1 and DBRX share?

1 benchmark has published results for both models. Claude 2.1 has 7 scored results on Noometry and DBRX has 21.

Related comparisons

Go deeper