Model comparison

Mixtral 8x22B vs Qwen-14B

Qwen-14B is the stronger model overall, scoring 31.4 to 27.1 on the Noometry Index.

Last verified . 12 shared benchmarks.

Mixtral 8x22B Mistral AI

27.1

Rank #333 Confirmed

Qwen-14B Alibaba (Qwen)

31.4

Rank #275 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Mixtral 8x22B scores higher in 5 categories and Qwen-14B in 2 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mixtral 8x22B leads 36.9 to 27.6.

Side by side

Mixtral 8x22B and Qwen-14B specifications
Mixtral 8x22BQwen-14B
ProviderMistral AIAlibaba (Qwen)
Noometry Index27.131.4
Released2024-04-172023-09-24
WeightsOpenOpen
Context window64K—
Max output64K—
Input $ / M tokens$2—
Output $ / M tokens$6—
Results tracked3418

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen-14B leads

Mixtral 8x22B: 24.2 (#329), Qwen-14B: 31.2 (#288)

Coding benchmarks
BenchmarkMixtral 8x22BQwen-14B
LMArena Coding11661071
WeirdML3.2%—
BigCodeBench Instruct40.6%—
BigCodeBench Complete50.2%—
HumanEval+72%—
MBPP+64.3%—

Agentic & Tool Use Not comparable

Mixtral 8x22B: 23.1 (#127), Qwen-14B: —

Agentic & Tool Use benchmarks
BenchmarkMixtral 8x22BQwen-14B
Cybench7.5%—

Reasoning Too close to call

Mixtral 8x22B: 19.9 (#248), Qwen-14B: 19.6 (#257)

Reasoning benchmarks
BenchmarkMixtral 8x22BQwen-14B
LMArena Hard Prompts11501027
Epoch Capabilities Index122.03113.03
DTBench55.1%—
BIG-Bench Hard—55%
ForecastBench56.3—
LAMBADA—71.1%
PIQA—79.9%

Math Qwen-14B leads

Mixtral 8x22B: 22.9 (#275), Qwen-14B: 31.2 (#227)

Math benchmarks
BenchmarkMixtral 8x22BQwen-14B
LMArena Math11841068
Omni-MATH16.3%—
MATH Level 524.2%—
GSM8K—61.3%

Knowledge Not comparable

Mixtral 8x22B: 15.1 (#293), Qwen-14B: —

Knowledge benchmarks
BenchmarkMixtral 8x22BQwen-14B
MMLU77.8%66.3%
GPQA Diamond34.1%—
MMLU-Pro46%—
GPQA (HELM)33.4%—
LMArena Expert1113—
ARC (AI2) Challenge—84.4%
BoolQ—86.2%

Multilingual Mixtral 8x22B leads

Mixtral 8x22B: 32.8 (#255), Qwen-14B: 27.5 (#275)

Multilingual benchmarks
BenchmarkMixtral 8x22BQwen-14B
LMArena Non-English11281041
LMArena Chinese11161077
LMArena French1166—
LMArena German1141—
LMArena Japanese1037—
LMArena Korean1057—
LMArena Russian1158—
LMArena Spanish1151—

Instruction Following Mixtral 8x22B leads

Mixtral 8x22B: 57.7 (#266), Qwen-14B: 52.4 (#289)

Instruction Following benchmarks
BenchmarkMixtral 8x22BQwen-14B
LMArena Instruction Following11471031
IFEval72.4%—

Long Context Mixtral 8x22B leads

Mixtral 8x22B: 34.7 (#247), Qwen-14B: 31.3 (#280)

Long Context benchmarks
BenchmarkMixtral 8x22BQwen-14B
LMArena Longer Query11441028

Writing & Preference Mixtral 8x22B leads

Mixtral 8x22B: 36.9 (#262), Qwen-14B: 27.6 (#299)

Writing & Preference benchmarks
BenchmarkMixtral 8x22BQwen-14B
LMArena Text11621051
LMArena Creative Writing11411028
LMArena Multi-Turn11301022
WildBench71.1%—

Frequently asked questions

Is Mixtral 8x22B better than Qwen-14B?

Qwen-14B is the stronger model overall, scoring 31.4 to 27.1 on the Noometry Index.

Is Mixtral 8x22B or Qwen-14B better for coding?

Qwen-14B scores higher on coding benchmarks: 31.2 versus 24.2 in the Noometry coding category.

How many benchmarks do Mixtral 8x22B and Qwen-14B share?

12 benchmarks have published results for both models. Mixtral 8x22B has 34 scored results on Noometry and Qwen-14B has 18.

Related comparisons

Go deeper