Model comparison

DBRX vs Qwen3 14B

Qwen3 14B is the stronger model overall, scoring 35.5 to 29.4 on the Noometry Index.

Last verified . 1 shared benchmarks.

DBRX Databricks

29.4

Rank #311 Confirmed

Qwen3 14B Alibaba (Qwen)

35.5

Rank #225 Confirmed

Summary

  • They share 1 benchmark with published results for both. DBRX scores higher in 1 category and Qwen3 14B in 4 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3 14B leads 39.3 to 14.9.
  • The biggest single-benchmark swing is GPQA Diamond: 32.9% for DBRX and 63.8% for Qwen3 14B.

Side by side

DBRX and Qwen3 14B specifications
DBRXQwen3 14B
ProviderDatabricksAlibaba (Qwen)
Noometry Index29.435.5
Released2024-03-272025-04
WeightsOpenOpen
Context window—131K
Max output—8K
Input $ / M tokens—$0.35
Output $ / M tokens—$1.40
Results tracked2112

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3 14B leads

DBRX: 32.9 (#266), Qwen3 14B: 37.3 (#195)

Coding benchmarks
BenchmarkDBRXQwen3 14B
SciCode—31.6%
LMArena Coding1132—
HumanEval+70.1%—
MBPP+55.8%—

Agentic & Tool Use Not comparable

DBRX: —, Qwen3 14B: 29.6 (#83)

Agentic & Tool Use benchmarks
BenchmarkDBRXQwen3 14B
Berkeley Function Calling Leaderboard—41%

Reasoning DBRX leads

DBRX: 21.4 (#222), Qwen3 14B: 18.5 (#280)

Reasoning benchmarks
BenchmarkDBRXQwen3 14B
Kagi LLM Benchmark—49.1%
CritPt—0%
Chess Puzzles—4%
LMArena Hard Prompts1113—
DTBench—64%
LMCA—18.2%
Epoch Capabilities Index—138.23

Math Qwen3 14B leads

DBRX: 24.3 (#269), Qwen3 14B: 38.6 (#133)

Math benchmarks
BenchmarkDBRXQwen3 14B
OTIS Mock AIME 2024-2025—66.4%
LMArena Math1145—
MATH Level 511.7%—

Knowledge Qwen3 14B leads

DBRX: 14.9 (#294), Qwen3 14B: 39.3 (#134)

Knowledge benchmarks
BenchmarkDBRXQwen3 14B
GPQA Diamond32.9%63.8%
Vectara Hallucination Rate—5.4%
LMArena Expert1076—

Multilingual Not comparable

DBRX: 29.3 (#268), Qwen3 14B: —

Multilingual benchmarks
BenchmarkDBRXQwen3 14B
LMArena Non-English1071—
LMArena Chinese1068—
LMArena French1096—
LMArena German1057—
LMArena Japanese990—
LMArena Korean993—
LMArena Russian1078—
LMArena Spanish1064—

Instruction Following Not comparable

DBRX: 57.5 (#270), Qwen3 14B: —

Instruction Following benchmarks
BenchmarkDBRXQwen3 14B
LMArena Instruction Following1112—

Long Context Qwen3 14B leads

DBRX: 33.7 (#258), Qwen3 14B: 38.1 (#204)

Long Context benchmarks
BenchmarkDBRXQwen3 14B
Fiction.LiveBench—62.5%
LMArena Longer Query1112—

Writing & Preference Not comparable

DBRX: 33.6 (#275), Qwen3 14B: —

Writing & Preference benchmarks
BenchmarkDBRXQwen3 14B
LMArena Text1119—
LMArena Creative Writing1104—
LMArena Multi-Turn1111—

Frequently asked questions

Is DBRX better than Qwen3 14B?

Qwen3 14B is the stronger model overall, scoring 35.5 to 29.4 on the Noometry Index.

Is DBRX or Qwen3 14B better for coding?

Qwen3 14B scores higher on coding benchmarks: 37.3 versus 32.9 in the Noometry coding category.

How many benchmarks do DBRX and Qwen3 14B share?

1 benchmark has published results for both models. DBRX has 21 scored results on Noometry and Qwen3 14B has 12.

Related comparisons

Go deeper