Model comparison

DBRX vs Llama 3.2 90B

DBRX is the stronger model overall, scoring 29.4 to 27.5 on the Noometry Index.

Last verified . 2 shared benchmarks.

DBRX Databricks

29.4

Rank #311 Confirmed

Llama 3.2 90B Meta

27.5

Rank #331 Confirmed

Summary

  • They share 2 benchmarks with published results for both. DBRX scores higher in 1 category and Llama 3.2 90B in 2 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in math, where DBRX leads 24.3 to 11.1.
  • The biggest single-benchmark swing is MATH Level 5: 11.7% for DBRX and 39.4% for Llama 3.2 90B.

Side by side

DBRX and Llama 3.2 90B specifications
DBRXLlama 3.2 90B
ProviderDatabricksMeta
Noometry Index29.427.5
Released2024-03-272024-09-24
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked219

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

DBRX: 32.9 (#266), Llama 3.2 90B: —

Coding benchmarks
BenchmarkDBRXLlama 3.2 90B
LMArena Coding1132—
HumanEval+70.1%—
MBPP+55.8%—

Agentic & Tool Use Not comparable

DBRX: —, Llama 3.2 90B: 30.0 (#80)

Agentic & Tool Use benchmarks
BenchmarkDBRXLlama 3.2 90B
BALROG—27.3%

Reasoning Too close to call

DBRX: 21.4 (#222), Llama 3.2 90B: 21.7 (#217)

Reasoning benchmarks
BenchmarkDBRXLlama 3.2 90B
EnigmaEval—0.4%
LMArena Hard Prompts1113—
Epoch Capabilities Index—125.5

Math DBRX leads

DBRX: 24.3 (#269), Llama 3.2 90B: 11.1 (#308)

Math benchmarks
BenchmarkDBRXLlama 3.2 90B
MATH Level 511.7%39.4%
OTIS Mock AIME 2024-2025—2.6%
LMArena Math1145—

Knowledge Llama 3.2 90B leads

DBRX: 14.9 (#294), Llama 3.2 90B: 21.7 (#274)

Knowledge benchmarks
BenchmarkDBRXLlama 3.2 90B
GPQA Diamond32.9%41%
LMArena Expert1076—
MMLU—80.3%

Multimodal Not comparable

DBRX: —, Llama 3.2 90B: 25.4 (#124)

Multimodal benchmarks
BenchmarkDBRXLlama 3.2 90B
LMArena Vision—1000
GeoBench—52%

Multilingual Not comparable

DBRX: 29.3 (#268), Llama 3.2 90B: —

Multilingual benchmarks
BenchmarkDBRXLlama 3.2 90B
LMArena Non-English1071—
LMArena Chinese1068—
LMArena French1096—
LMArena German1057—
LMArena Japanese990—
LMArena Korean993—
LMArena Russian1078—
LMArena Spanish1064—

Instruction Following Not comparable

DBRX: 57.5 (#270), Llama 3.2 90B: —

Instruction Following benchmarks
BenchmarkDBRXLlama 3.2 90B
LMArena Instruction Following1112—

Long Context Not comparable

DBRX: 33.7 (#258), Llama 3.2 90B: —

Long Context benchmarks
BenchmarkDBRXLlama 3.2 90B
LMArena Longer Query1112—

Writing & Preference Not comparable

DBRX: 33.6 (#275), Llama 3.2 90B: —

Writing & Preference benchmarks
BenchmarkDBRXLlama 3.2 90B
LMArena Text1119—
LMArena Creative Writing1104—
LMArena Multi-Turn1111—

Frequently asked questions

Is DBRX better than Llama 3.2 90B?

DBRX is the stronger model overall, scoring 29.4 to 27.5 on the Noometry Index.

How many benchmarks do DBRX and Llama 3.2 90B share?

2 benchmarks have published results for both models. DBRX has 21 scored results on Noometry and Llama 3.2 90B has 9.

Related comparisons

Go deeper