Model comparison

DBRX vs Qwen3-4B

Qwen3-4B is the stronger model overall, scoring 31.9 to 29.4 on the Noometry Index.

Last verified . 1 shared benchmarks.

DBRX Databricks

29.4

Rank #311 Confirmed

Qwen3-4B Alibaba (Qwen)

31.9

Rank #264 Confirmed

Summary

  • They share 1 benchmark with published results for both. DBRX scores higher in 1 category and Qwen3-4B in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3-4B leads 33.0 to 14.9.
  • The biggest single-benchmark swing is GPQA Diamond: 32.9% for DBRX and 52.3% for Qwen3-4B.

Side by side

DBRX and Qwen3-4B specifications
DBRXQwen3-4B
ProviderDatabricksAlibaba (Qwen)
Noometry Index29.431.9
Released2024-03-272025-04-29
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked216

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

DBRX: 32.9 (#266), Qwen3-4B: —

Coding benchmarks
BenchmarkDBRXQwen3-4B
LMArena Coding1132—
HumanEval+70.1%—
MBPP+55.8%—

Agentic & Tool Use Not comparable

DBRX: —, Qwen3-4B: 27.6 (#100)

Agentic & Tool Use benchmarks
BenchmarkDBRXQwen3-4B
Berkeley Function Calling Leaderboard—35.7%

Reasoning DBRX leads

DBRX: 21.4 (#222), Qwen3-4B: 19.2 (#268)

Reasoning benchmarks
BenchmarkDBRXQwen3-4B
Chess Puzzles—4%
LMArena Hard Prompts1113—

Math Qwen3-4B leads

DBRX: 24.3 (#269), Qwen3-4B: 29.7 (#240)

Math benchmarks
BenchmarkDBRXQwen3-4B
MathArena Final-Answer Competitions—38.5%
OTIS Mock AIME 2024-2025—52.2%
LMArena Math1145—
MATH Level 511.7%—

Knowledge Qwen3-4B leads

DBRX: 14.9 (#294), Qwen3-4B: 33.0 (#208)

Knowledge benchmarks
BenchmarkDBRXQwen3-4B
GPQA Diamond32.9%52.3%
Vectara Hallucination Rate—5.7%
LMArena Expert1076—

Multilingual Not comparable

DBRX: 29.3 (#268), Qwen3-4B: —

Multilingual benchmarks
BenchmarkDBRXQwen3-4B
LMArena Non-English1071—
LMArena Chinese1068—
LMArena French1096—
LMArena German1057—
LMArena Japanese990—
LMArena Korean993—
LMArena Russian1078—
LMArena Spanish1064—

Instruction Following Not comparable

DBRX: 57.5 (#270), Qwen3-4B: —

Instruction Following benchmarks
BenchmarkDBRXQwen3-4B
LMArena Instruction Following1112—

Long Context Not comparable

DBRX: 33.7 (#258), Qwen3-4B: —

Long Context benchmarks
BenchmarkDBRXQwen3-4B
LMArena Longer Query1112—

Writing & Preference Not comparable

DBRX: 33.6 (#275), Qwen3-4B: —

Writing & Preference benchmarks
BenchmarkDBRXQwen3-4B
LMArena Text1119—
LMArena Creative Writing1104—
LMArena Multi-Turn1111—

Frequently asked questions

Is DBRX better than Qwen3-4B?

Qwen3-4B is the stronger model overall, scoring 31.9 to 29.4 on the Noometry Index.

How many benchmarks do DBRX and Qwen3-4B share?

1 benchmark has published results for both models. DBRX has 21 scored results on Noometry and Qwen3-4B has 6.

Related comparisons

Go deeper