Model comparison

Mixtral 8x7B vs Qwen1.5 4b Chat

Qwen1.5 4b Chat is the stronger model overall, scoring 28.8 to 27.1 on the Noometry Index.

Last verified . 13 shared benchmarks.

Mixtral 8x7B Mistral AI

27.1

Rank #334 Confirmed

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Mixtral 8x7B scores higher in 5 categories and Qwen1.5 4b Chat in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen1.5 4b Chat leads 26.7 to 11.0.

Side by side

Mixtral 8x7B and Qwen1.5 4b Chat specifications
Mixtral 8x7BQwen1.5 4b Chat
ProviderMistral AIAlibaba (Qwen)
Noometry Index27.128.8
Released2023-12-11—
WeightsOpenOpen
Context window32K—
Max output32K—
Input $ / M tokens$0.70—
Output $ / M tokens$0.70—
Results tracked3813

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mixtral 8x7B leads

Mixtral 8x7B: 32.8 (#269), Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkMixtral 8x7BQwen1.5 4b Chat
LMArena Coding1126999
HumanEval+39.6%—
MBPP+49.7%—

Reasoning Too close to call

Mixtral 8x7B: 18.2 (#285), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkMixtral 8x7BQwen1.5 4b Chat
LMArena Hard Prompts1115976
DTBench49.6%—
Adversarial NLI55.2%—
Epoch Capabilities Index118.47—
ForecastBench56.3—
HellaSwag86.7%—
PIQA83.6%—
WinoGrande77.2%—

Math Qwen1.5 4b Chat leads

Mixtral 8x7B: 18.8 (#289), Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkMixtral 8x7BQwen1.5 4b Chat
LMArena Math11471026
Omni-MATH10.5%—
MATH Level 510%—
GSM8K74.4%—

Knowledge Qwen1.5 4b Chat leads

Mixtral 8x7B: 11.0 (#301), Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkMixtral 8x7BQwen1.5 4b Chat
LMArena Expert1088980
GPQA Diamond30.6%—
MMLU-Pro33.5%—
GPQA (HELM)29.6%—
ARC (AI2) Challenge87.3%—
MMLU70.6%—
OpenBookQA85.8%—
TriviaQA82.2%—

Multilingual Mixtral 8x7B leads

Mixtral 8x7B: 29.6 (#266), Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkMixtral 8x7BQwen1.5 4b Chat
LMArena Non-English1077979
LMArena Chinese10551024
LMArena German1114902
LMArena Russian1090952
LMArena French1166—
LMArena Japanese931—
LMArena Korean968—
LMArena Spanish1111—

Instruction Following Mixtral 8x7B leads

Mixtral 8x7B: 51.0 (#297), Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkMixtral 8x7BQwen1.5 4b Chat
LMArena Instruction Following1109978
IFEval57.5%—

Long Context Mixtral 8x7B leads

Mixtral 8x7B: 33.4 (#260), Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkMixtral 8x7BQwen1.5 4b Chat
LMArena Longer Query1103988

Writing & Preference Mixtral 8x7B leads

Mixtral 8x7B: 34.2 (#270), Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkMixtral 8x7BQwen1.5 4b Chat
LMArena Text1132997
LMArena Creative Writing1109969
LMArena Multi-Turn1115977
WildBench67.3%—

Frequently asked questions

Is Mixtral 8x7B better than Qwen1.5 4b Chat?

Qwen1.5 4b Chat is the stronger model overall, scoring 28.8 to 27.1 on the Noometry Index.

Is Mixtral 8x7B or Qwen1.5 4b Chat better for coding?

Mixtral 8x7B scores higher on coding benchmarks: 32.8 versus 29.1 in the Noometry coding category.

How many benchmarks do Mixtral 8x7B and Qwen1.5 4b Chat share?

13 benchmarks have published results for both models. Mixtral 8x7B has 38 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper