Model comparison

DBRX vs Mistral Small 3.1

Mistral Small 3.1 is the stronger model overall, scoring 31.7 to 29.4 on the Noometry Index.

Last verified . 18 shared benchmarks.

DBRX Databricks

29.4

Rank #311 Confirmed

Mistral Small 3.1 Mistral AI

31.7

Rank #269 Confirmed

Summary

  • They share 18 benchmarks with published results for both. DBRX scores higher in 2 categories and Mistral Small 3.1 in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Mistral Small 3.1 leads 41.2 to 29.3.
  • The biggest single-benchmark swing is GPQA Diamond: 32.9% for DBRX and 41.9% for Mistral Small 3.1.

Side by side

DBRX and Mistral Small 3.1 specifications
DBRXMistral Small 3.1
ProviderDatabricksMistral AI
Noometry Index29.431.7
Released2024-03-272025-03-17
WeightsOpenOpen
Context window—128K
Max output—102K
Input $ / M tokens—$0.35
Output $ / M tokens—$0.56
Results tracked2128

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Small 3.1 leads

DBRX: 32.9 (#266), Mistral Small 3.1: 38.3 (#179)

Coding benchmarks
BenchmarkDBRXMistral Small 3.1
LMArena Coding11321309
HumanEval+70.1%—
MBPP+55.8%—

Reasoning DBRX leads

DBRX: 21.4 (#222), Mistral Small 3.1: 19.7 (#254)

Reasoning benchmarks
BenchmarkDBRXMistral Small 3.1
LMArena Hard Prompts11131278
Chess Puzzles—1%
Epoch Capabilities Index—127.48

Math DBRX leads

DBRX: 24.3 (#269), Mistral Small 3.1: 14.7 (#301)

Math benchmarks
BenchmarkDBRXMistral Small 3.1
LMArena Math11451262
OTIS Mock AIME 2024-2025—3.9%
Omni-MATH—24.8%
MATH Level 511.7%—

Knowledge Mistral Small 3.1 leads

DBRX: 14.9 (#294), Mistral Small 3.1: 22.6 (#271)

Knowledge benchmarks
BenchmarkDBRXMistral Small 3.1
GPQA Diamond32.9%41.9%
LMArena Expert10761257
MMLU-Pro—61%
GPQA (HELM)—39.2%

Multimodal Not comparable

DBRX: —, Mistral Small 3.1: 33.2 (#99)

Multimodal benchmarks
BenchmarkDBRXMistral Small 3.1
LMArena Vision—1136

Multilingual Mistral Small 3.1 leads

DBRX: 29.3 (#268), Mistral Small 3.1: 41.2 (#209)

Multilingual benchmarks
BenchmarkDBRXMistral Small 3.1
LMArena Non-English10711255
LMArena Chinese10681253
LMArena French10961273
LMArena German10571266
LMArena Japanese9901208
LMArena Korean9931206
LMArena Russian10781263
LMArena Spanish10641283

Instruction Following Mistral Small 3.1 leads

DBRX: 57.5 (#270), Mistral Small 3.1: 63.6 (#230)

Instruction Following benchmarks
BenchmarkDBRXMistral Small 3.1
LMArena Instruction Following11121264
IFEval—75%

Long Context Mistral Small 3.1 leads

DBRX: 33.7 (#258), Mistral Small 3.1: 39.5 (#178)

Long Context benchmarks
BenchmarkDBRXMistral Small 3.1
LMArena Longer Query11121299

Writing & Preference Mistral Small 3.1 leads

DBRX: 33.6 (#275), Mistral Small 3.1: 37.0 (#259)

Writing & Preference benchmarks
BenchmarkDBRXMistral Small 3.1
LMArena Text11191277
LMArena Creative Writing11041253
LMArena Multi-Turn11111270
EQ-Bench Creative Writing—761
WildBench—78.8%

Frequently asked questions

Is DBRX better than Mistral Small 3.1?

Mistral Small 3.1 is the stronger model overall, scoring 31.7 to 29.4 on the Noometry Index.

Is DBRX or Mistral Small 3.1 better for coding?

Mistral Small 3.1 scores higher on coding benchmarks: 38.3 versus 32.9 in the Noometry coding category.

How many benchmarks do DBRX and Mistral Small 3.1 share?

18 benchmarks have published results for both models. DBRX has 21 scored results on Noometry and Mistral Small 3.1 has 28.

Related comparisons

Go deeper