Model comparison

DBRX vs DeepSeek V4.1 Flash

DeepSeek V4.1 Flash is the stronger model overall, scoring 52.8 to 29.4 on the Noometry Index.

Last verified . 18 shared benchmarks.

DBRX Databricks

29.4

Rank #311 Confirmed

DeepSeek V4.1 Flash DeepSeek

52.8

Rank #38 Confirmed

Summary

  • They share 18 benchmarks with published results for both. DBRX scores higher in 0 categories and DeepSeek V4.1 Flash in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where DeepSeek V4.1 Flash leads 57.9 to 14.9.
  • The biggest single-benchmark swing is GPQA Diamond: 32.9% for DBRX and 89.8% for DeepSeek V4.1 Flash.

Side by side

DBRX and DeepSeek V4.1 Flash specifications
DBRXDeepSeek V4.1 Flash
ProviderDatabricksDeepSeek
Noometry Index29.452.8
Released2024-03-272026-09-09
WeightsOpenOpen
Context window—1M
Max output—393K
Input $ / M tokens—$0.15
Output $ / M tokens—$0.60
Results tracked2137

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek V4.1 Flash leads

DBRX: 32.9 (#266), DeepSeek V4.1 Flash: 52.9 (#32)

Coding benchmarks
BenchmarkDBRXDeepSeek V4.1 Flash
LMArena Coding11321506
LMArena WebDev—1619
SciCode—51.9%
ALE-Bench—1,092
HumanEval+70.1%—
MBPP+55.8%—

Agentic & Tool Use Not comparable

DBRX: —, DeepSeek V4.1 Flash: 31.2 (#69)

Agentic & Tool Use benchmarks
BenchmarkDBRXDeepSeek V4.1 Flash
APEX-Agents—39.5%
GDP.pdf—19.8%

Reasoning DeepSeek V4.1 Flash leads

DBRX: 21.4 (#222), DeepSeek V4.1 Flash: 50.2 (#36)

Reasoning benchmarks
BenchmarkDBRXDeepSeek V4.1 Flash
LMArena Hard Prompts11131483
NYT Connections (extended)—89.6%
CritPt—14.3%
Mystery Game Puzzles—43%
DTBench—89.9%
LMCA—47%
Surface Evolver Bench—46.3%
Epoch Capabilities Index—154.9

Math DeepSeek V4.1 Flash leads

DBRX: 24.3 (#269), DeepSeek V4.1 Flash: 66.7 (#25)

Math benchmarks
BenchmarkDBRXDeepSeek V4.1 Flash
LMArena Math11451477
FrontierMath (Tiers 1-3)—67.4%
FrontierMath Tier 4—26.8%
OTIS Mock AIME 2024-2025—98.3%
ProofBench—54%
MATH Level 511.7%—

Knowledge DeepSeek V4.1 Flash leads

DBRX: 14.9 (#294), DeepSeek V4.1 Flash: 57.9 (#38)

Knowledge benchmarks
BenchmarkDBRXDeepSeek V4.1 Flash
GPQA Diamond32.9%89.8%
LMArena Expert10761506

Multimodal Not comparable

DBRX: —, DeepSeek V4.1 Flash: 39.1 (#61)

Multimodal benchmarks
BenchmarkDBRXDeepSeek V4.1 Flash
LMArena Vision—1277
Furniture Assembly—34.2%

Multilingual DeepSeek V4.1 Flash leads

DBRX: 29.3 (#268), DeepSeek V4.1 Flash: 55.0 (#35)

Multilingual benchmarks
BenchmarkDBRXDeepSeek V4.1 Flash
LMArena Non-English10711448
LMArena Chinese10681497
LMArena French10961452
LMArena German10571484
LMArena Japanese9901412
LMArena Korean9931452
LMArena Russian10781471
LMArena Spanish10641459

Instruction Following DeepSeek V4.1 Flash leads

DBRX: 57.5 (#270), DeepSeek V4.1 Flash: 77.3 (#26)

Instruction Following benchmarks
BenchmarkDBRXDeepSeek V4.1 Flash
LMArena Instruction Following11121474

Long Context DeepSeek V4.1 Flash leads

DBRX: 33.7 (#258), DeepSeek V4.1 Flash: 45.2 (#47)

Long Context benchmarks
BenchmarkDBRXDeepSeek V4.1 Flash
LMArena Longer Query11121475

Writing & Preference DeepSeek V4.1 Flash leads

DBRX: 33.6 (#275), DeepSeek V4.1 Flash: 65.4 (#48)

Writing & Preference benchmarks
BenchmarkDBRXDeepSeek V4.1 Flash
LMArena Text11191462
LMArena Creative Writing11041435
LMArena Multi-Turn11111457
EQ-Bench Creative Writing—1540

Frequently asked questions

Is DBRX better than DeepSeek V4.1 Flash?

DeepSeek V4.1 Flash is the stronger model overall, scoring 52.8 to 29.4 on the Noometry Index.

Is DBRX or DeepSeek V4.1 Flash better for coding?

DeepSeek V4.1 Flash scores higher on coding benchmarks: 52.9 versus 32.9 in the Noometry coding category.

How many benchmarks do DBRX and DeepSeek V4.1 Flash share?

18 benchmarks have published results for both models. DBRX has 21 scored results on Noometry and DeepSeek V4.1 Flash has 37.

Related comparisons

Go deeper