Model comparison

Mixtral 8x22B vs Qwen1.5 4b Chat

Qwen1.5 4b Chat is the stronger model overall, scoring 28.8 to 27.1 on the Noometry Index.

Last verified . 13 shared benchmarks.

Mixtral 8x22B Mistral AI

27.1

Rank #333 Confirmed

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Mixtral 8x22B scores higher in 5 categories and Qwen1.5 4b Chat in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mixtral 8x22B leads 36.9 to 23.8.

Side by side

Mixtral 8x22B and Qwen1.5 4b Chat specifications
Mixtral 8x22BQwen1.5 4b Chat
ProviderMistral AIAlibaba (Qwen)
Noometry Index27.128.8
Released2024-04-17—
WeightsOpenOpen
Context window64K—
Max output64K—
Input $ / M tokens$2—
Output $ / M tokens$6—
Results tracked3413

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen1.5 4b Chat leads

Mixtral 8x22B: 24.2 (#329), Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkMixtral 8x22BQwen1.5 4b Chat
LMArena Coding1166999
WeirdML3.2%—
BigCodeBench Instruct40.6%—
BigCodeBench Complete50.2%—
HumanEval+72%—
MBPP+64.3%—

Agentic & Tool Use Not comparable

Mixtral 8x22B: 23.1 (#127), Qwen1.5 4b Chat: —

Agentic & Tool Use benchmarks
BenchmarkMixtral 8x22BQwen1.5 4b Chat
Cybench7.5%—

Reasoning Mixtral 8x22B leads

Mixtral 8x22B: 19.9 (#248), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkMixtral 8x22BQwen1.5 4b Chat
LMArena Hard Prompts1150976
DTBench55.1%—
Epoch Capabilities Index122.03—
ForecastBench56.3—

Math Qwen1.5 4b Chat leads

Mixtral 8x22B: 22.9 (#275), Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkMixtral 8x22BQwen1.5 4b Chat
LMArena Math11841026
Omni-MATH16.3%—
MATH Level 524.2%—

Knowledge Qwen1.5 4b Chat leads

Mixtral 8x22B: 15.1 (#293), Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkMixtral 8x22BQwen1.5 4b Chat
LMArena Expert1113980
GPQA Diamond34.1%—
MMLU-Pro46%—
GPQA (HELM)33.4%—
MMLU77.8%—

Multilingual Mixtral 8x22B leads

Mixtral 8x22B: 32.8 (#255), Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkMixtral 8x22BQwen1.5 4b Chat
LMArena Non-English1128979
LMArena Chinese11161024
LMArena German1141902
LMArena Russian1158952
LMArena French1166—
LMArena Japanese1037—
LMArena Korean1057—
LMArena Spanish1151—

Instruction Following Mixtral 8x22B leads

Mixtral 8x22B: 57.7 (#266), Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkMixtral 8x22BQwen1.5 4b Chat
LMArena Instruction Following1147978
IFEval72.4%—

Long Context Mixtral 8x22B leads

Mixtral 8x22B: 34.7 (#247), Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkMixtral 8x22BQwen1.5 4b Chat
LMArena Longer Query1144988

Writing & Preference Mixtral 8x22B leads

Mixtral 8x22B: 36.9 (#262), Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkMixtral 8x22BQwen1.5 4b Chat
LMArena Text1162997
LMArena Creative Writing1141969
LMArena Multi-Turn1130977
WildBench71.1%—

Frequently asked questions

Is Mixtral 8x22B better than Qwen1.5 4b Chat?

Qwen1.5 4b Chat is the stronger model overall, scoring 28.8 to 27.1 on the Noometry Index.

Is Mixtral 8x22B or Qwen1.5 4b Chat better for coding?

Qwen1.5 4b Chat scores higher on coding benchmarks: 29.1 versus 24.2 in the Noometry coding category.

How many benchmarks do Mixtral 8x22B and Qwen1.5 4b Chat share?

13 benchmarks have published results for both models. Mixtral 8x22B has 34 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper