Model comparison

Mixtral 8x7B vs Qwen-14B

Qwen-14B is the stronger model overall, scoring 31.4 to 27.1 on the Noometry Index.

Last verified . 15 shared benchmarks.

Mixtral 8x7B Mistral AI

27.1

Rank #334 Confirmed

Qwen-14B Alibaba (Qwen)

31.4

Rank #275 Confirmed

Summary

  • They share 15 benchmarks with published results for both. Mixtral 8x7B scores higher in 4 categories and Qwen-14B in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen-14B leads 31.2 to 18.8.

Side by side

Mixtral 8x7B and Qwen-14B specifications
Mixtral 8x7BQwen-14B
ProviderMistral AIAlibaba (Qwen)
Noometry Index27.131.4
Released2023-12-112023-09-24
WeightsOpenOpen
Context window32K—
Max output32K—
Input $ / M tokens$0.70—
Output $ / M tokens$0.70—
Results tracked3818

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mixtral 8x7B leads

Mixtral 8x7B: 32.8 (#269), Qwen-14B: 31.2 (#288)

Coding benchmarks
BenchmarkMixtral 8x7BQwen-14B
LMArena Coding11261071
HumanEval+39.6%—
MBPP+49.7%—

Reasoning Qwen-14B leads

Mixtral 8x7B: 18.2 (#285), Qwen-14B: 19.6 (#257)

Reasoning benchmarks
BenchmarkMixtral 8x7BQwen-14B
LMArena Hard Prompts11151027
Epoch Capabilities Index118.47113.03
PIQA83.6%79.9%
DTBench49.6%—
Adversarial NLI55.2%—
BIG-Bench Hard—55%
ForecastBench56.3—
HellaSwag86.7%—
LAMBADA—71.1%
WinoGrande77.2%—

Math Qwen-14B leads

Mixtral 8x7B: 18.8 (#289), Qwen-14B: 31.2 (#227)

Math benchmarks
BenchmarkMixtral 8x7BQwen-14B
LMArena Math11471068
GSM8K74.4%61.3%
Omni-MATH10.5%—
MATH Level 510%—

Knowledge Not comparable

Mixtral 8x7B: 11.0 (#301), Qwen-14B: —

Knowledge benchmarks
BenchmarkMixtral 8x7BQwen-14B
ARC (AI2) Challenge87.3%84.4%
MMLU70.6%66.3%
GPQA Diamond30.6%—
MMLU-Pro33.5%—
GPQA (HELM)29.6%—
LMArena Expert1088—
BoolQ—86.2%
OpenBookQA85.8%—
TriviaQA82.2%—

Multilingual Mixtral 8x7B leads

Mixtral 8x7B: 29.6 (#266), Qwen-14B: 27.5 (#275)

Multilingual benchmarks
BenchmarkMixtral 8x7BQwen-14B
LMArena Non-English10771041
LMArena Chinese10551077
LMArena French1166—
LMArena German1114—
LMArena Japanese931—
LMArena Korean968—
LMArena Russian1090—
LMArena Spanish1111—

Instruction Following Qwen-14B leads

Mixtral 8x7B: 51.0 (#297), Qwen-14B: 52.4 (#289)

Instruction Following benchmarks
BenchmarkMixtral 8x7BQwen-14B
LMArena Instruction Following11091031
IFEval57.5%—

Long Context Mixtral 8x7B leads

Mixtral 8x7B: 33.4 (#260), Qwen-14B: 31.3 (#280)

Long Context benchmarks
BenchmarkMixtral 8x7BQwen-14B
LMArena Longer Query11031028

Writing & Preference Mixtral 8x7B leads

Mixtral 8x7B: 34.2 (#270), Qwen-14B: 27.6 (#299)

Writing & Preference benchmarks
BenchmarkMixtral 8x7BQwen-14B
LMArena Text11321051
LMArena Creative Writing11091028
LMArena Multi-Turn11151022
WildBench67.3%—

Frequently asked questions

Is Mixtral 8x7B better than Qwen-14B?

Qwen-14B is the stronger model overall, scoring 31.4 to 27.1 on the Noometry Index.

Is Mixtral 8x7B or Qwen-14B better for coding?

Mixtral 8x7B scores higher on coding benchmarks: 32.8 versus 31.2 in the Noometry coding category.

How many benchmarks do Mixtral 8x7B and Qwen-14B share?

15 benchmarks have published results for both models. Mixtral 8x7B has 38 scored results on Noometry and Qwen-14B has 18.

Related comparisons

Go deeper