Model comparison

DBRX vs Llama-3.3-70B-Instruct

Llama-3.3-70B-Instruct is the stronger model overall, scoring 30.6 to 29.4 on the Noometry Index.

Last verified . 19 shared benchmarks.

DBRX Databricks

29.4

Rank #311 Confirmed

Llama-3.3-70B-Instruct Meta

30.6

Rank #291 Confirmed

Summary

  • They share 19 benchmarks with published results for both. DBRX scores higher in 4 categories and Llama-3.3-70B-Instruct in 4 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Llama-3.3-70B-Instruct leads 30.6 to 14.9.
  • The biggest single-benchmark swing is MATH Level 5: 11.7% for DBRX and 41.6% for Llama-3.3-70B-Instruct.

Side by side

DBRX and Llama-3.3-70B-Instruct specifications
DBRXLlama-3.3-70B-Instruct
ProviderDatabricksMeta
Noometry Index29.430.6
Released2024-03-272024-12-06
WeightsOpenOpen
Context window—128K
Max output—4K
Input $ / M tokens—$0.10
Output $ / M tokens—$0.32
Results tracked2143

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DBRX leads

DBRX: 32.9 (#266), Llama-3.3-70B-Instruct: 31.0 (#290)

Coding benchmarks
BenchmarkDBRXLlama-3.3-70B-Instruct
LMArena Coding11321268
SciCode—26%
WeirdML—14.4%
BigCodeBench Instruct—46.9%
LiveBench Coding—36.6%
BigCodeBench Complete—57.5%
HumanEval+70.1%—
MBPP+55.8%—

Agentic & Tool Use Not comparable

DBRX: —, Llama-3.3-70B-Instruct: 25.8 (#105)

Agentic & Tool Use benchmarks
BenchmarkDBRXLlama-3.3-70B-Instruct
Berkeley Function Calling Leaderboard—31.9%
BALROG—23%

Reasoning DBRX leads

DBRX: 21.4 (#222), Llama-3.3-70B-Instruct: 14.1 (#327)

Reasoning benchmarks
BenchmarkDBRXLlama-3.3-70B-Instruct
LMArena Hard Prompts11131257
SimpleBench—19.9%
CritPt—0%
LiveBench Reasoning—50.8%
DTBench—59.5%
LiveBench Data Analysis—49.5%
LMCA—17.5%
Epoch Capabilities Index—127.33
ForecastBench—58.6
LiveBench—50.2%

Math DBRX leads

DBRX: 24.3 (#269), Llama-3.3-70B-Instruct: 15.3 (#298)

Math benchmarks
BenchmarkDBRXLlama-3.3-70B-Instruct
LMArena Math11451267
MATH Level 511.7%41.6%
OTIS Mock AIME 2024-2025—5.1%
LiveBench Math—42.2%

Knowledge Llama-3.3-70B-Instruct leads

DBRX: 14.9 (#294), Llama-3.3-70B-Instruct: 30.6 (#226)

Knowledge benchmarks
BenchmarkDBRXLlama-3.3-70B-Instruct
GPQA Diamond32.9%47.4%
LMArena Expert10761225
Confabulations—22.8%
Vectara Hallucination Rate—4.1%
MMLU—86.3%

Multilingual Llama-3.3-70B-Instruct leads

DBRX: 29.3 (#268), Llama-3.3-70B-Instruct: 39.9 (#220)

Multilingual benchmarks
BenchmarkDBRXLlama-3.3-70B-Instruct
LMArena Non-English10711236
LMArena Chinese10681217
LMArena French10961281
LMArena German10571251
LMArena Japanese9901150
LMArena Korean9931143
LMArena Russian10781252
LMArena Spanish10641270

Instruction Following Llama-3.3-70B-Instruct leads

DBRX: 57.5 (#270), Llama-3.3-70B-Instruct: 71.1 (#157)

Instruction Following benchmarks
BenchmarkDBRXLlama-3.3-70B-Instruct
LMArena Instruction Following11121242
LiveBench Instruction Following—82.7%

Long Context DBRX leads

DBRX: 33.7 (#258), Llama-3.3-70B-Instruct: 26.4 (#295)

Long Context benchmarks
BenchmarkDBRXLlama-3.3-70B-Instruct
LMArena Longer Query11121256
Fiction.LiveBench—33.3%

Writing & Preference Llama-3.3-70B-Instruct leads

DBRX: 33.6 (#275), Llama-3.3-70B-Instruct: 47.6 (#207)

Writing & Preference benchmarks
BenchmarkDBRXLlama-3.3-70B-Instruct
LMArena Text11191274
LMArena Creative Writing11041250
LMArena Multi-Turn11111280
LiveBench Language—39.2%

Frequently asked questions

Is DBRX better than Llama-3.3-70B-Instruct?

Llama-3.3-70B-Instruct is the stronger model overall, scoring 30.6 to 29.4 on the Noometry Index.

Is DBRX or Llama-3.3-70B-Instruct better for coding?

DBRX scores higher on coding benchmarks: 32.9 versus 31.0 in the Noometry coding category.

How many benchmarks do DBRX and Llama-3.3-70B-Instruct share?

19 benchmarks have published results for both models. DBRX has 21 scored results on Noometry and Llama-3.3-70B-Instruct has 43.

Related comparisons

Go deeper