Model comparison

DBRX vs Llama 3.2 1B

DBRX is the stronger model overall, scoring 29.4 to 20.1 on the Noometry Index.

Last verified . 14 shared benchmarks.

DBRX Databricks

29.4

Rank #311 Confirmed

Llama 3.2 1B Meta

20.1

Rank #354 Confirmed

Summary

  • They share 14 benchmarks with published results for both. DBRX scores higher in 8 categories and Llama 3.2 1B in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where DBRX leads 24.3 to 10.4.
  • The biggest single-benchmark swing is GPQA Diamond: 32.9% for DBRX and 23.9% for Llama 3.2 1B.

Side by side

DBRX and Llama 3.2 1B specifications
DBRXLlama 3.2 1B
ProviderDatabricksMeta
Noometry Index29.420.1
Released2024-03-272024-09-24
WeightsOpenOpen
Context window—60K
Max output—54K
Input $ / M tokens—$0.027
Output $ / M tokens—$0.20
Results tracked2122

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DBRX leads

DBRX: 32.9 (#266), Llama 3.2 1B: 21.1 (#338)

Coding benchmarks
BenchmarkDBRXLlama 3.2 1B
LMArena Coding11321070
BigCodeBench Instruct—8.2%
BigCodeBench Complete—11.3%
HumanEval+70.1%—
MBPP+55.8%—

Agentic & Tool Use Not comparable

DBRX: —, Llama 3.2 1B: 14.6 (#150)

Agentic & Tool Use benchmarks
BenchmarkDBRXLlama 3.2 1B
Berkeley Function Calling Leaderboard—10.8%
BALROG—6.6%

Reasoning DBRX leads

DBRX: 21.4 (#222), Llama 3.2 1B: 16.2 (#308)

Reasoning benchmarks
BenchmarkDBRXLlama 3.2 1B
LMArena Hard Prompts11131044
Chess Puzzles—0%
Epoch Capabilities Index—101.99

Math DBRX leads

DBRX: 24.3 (#269), Llama 3.2 1B: 10.4 (#313)

Math benchmarks
BenchmarkDBRXLlama 3.2 1B
LMArena Math11451086
OTIS Mock AIME 2024-2025—0.6%
MATH Level 511.7%—

Knowledge DBRX leads

DBRX: 14.9 (#294), Llama 3.2 1B: 7.2 (#312)

Knowledge benchmarks
BenchmarkDBRXLlama 3.2 1B
GPQA Diamond32.9%23.9%
LMArena Expert10761007

Multilingual DBRX leads

DBRX: 29.3 (#268), Llama 3.2 1B: 23.8 (#292)

Multilingual benchmarks
BenchmarkDBRXLlama 3.2 1B
LMArena Non-English1071973
LMArena Chinese1068959
LMArena German10571014
LMArena Russian1078941
LMArena French1096—
LMArena Japanese990—
LMArena Korean993—
LMArena Spanish1064—

Instruction Following DBRX leads

DBRX: 57.5 (#270), Llama 3.2 1B: 52.4 (#290)

Instruction Following benchmarks
BenchmarkDBRXLlama 3.2 1B
LMArena Instruction Following11121031

Long Context DBRX leads

DBRX: 33.7 (#258), Llama 3.2 1B: 31.9 (#274)

Long Context benchmarks
BenchmarkDBRXLlama 3.2 1B
LMArena Longer Query11121050

Writing & Preference DBRX leads

DBRX: 33.6 (#275), Llama 3.2 1B: 21.3 (#310)

Writing & Preference benchmarks
BenchmarkDBRXLlama 3.2 1B
LMArena Text11191055
LMArena Creative Writing11041033
LMArena Multi-Turn11111030
EQ-Bench Creative Writing—200

Frequently asked questions

Is DBRX better than Llama 3.2 1B?

DBRX is the stronger model overall, scoring 29.4 to 20.1 on the Noometry Index.

Is DBRX or Llama 3.2 1B better for coding?

DBRX scores higher on coding benchmarks: 32.9 versus 21.1 in the Noometry coding category.

How many benchmarks do DBRX and Llama 3.2 1B share?

14 benchmarks have published results for both models. DBRX has 21 scored results on Noometry and Llama 3.2 1B has 22.

Related comparisons

Go deeper