Model comparison

Mixtral 8x22B vs Qwen3.5 Max Preview

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 27.1 on the Noometry Index.

Last verified . 17 shared benchmarks.

Mixtral 8x22B Mistral AI

27.1

Rank #333 Confirmed

Qwen3.5 Max Preview Alibaba (Qwen)

45.3

Rank #71 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Mixtral 8x22B scores higher in 0 categories and Qwen3.5 Max Preview in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen3.5 Max Preview leads 66.0 to 36.9.
  • Mixtral 8x22B has downloadable open weights; the other is API-only.

Side by side

Mixtral 8x22B and Qwen3.5 Max Preview specifications
Mixtral 8x22BQwen3.5 Max Preview
ProviderMistral AIAlibaba (Qwen)
Noometry Index27.145.3
Released2024-04-17—
WeightsOpenProprietary
Context window64K—
Max output64K—
Input $ / M tokens$2—
Output $ / M tokens$6—
Results tracked3417

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.5 Max Preview leads

Mixtral 8x22B: 24.2 (#329), Qwen3.5 Max Preview: 44.0 (#77)

Coding benchmarks
BenchmarkMixtral 8x22BQwen3.5 Max Preview
LMArena Coding11661487
WeirdML3.2%—
BigCodeBench Instruct40.6%—
BigCodeBench Complete50.2%—
HumanEval+72%—
MBPP+64.3%—

Agentic & Tool Use Not comparable

Mixtral 8x22B: 23.1 (#127), Qwen3.5 Max Preview: —

Agentic & Tool Use benchmarks
BenchmarkMixtral 8x22BQwen3.5 Max Preview
Cybench7.5%—

Reasoning Qwen3.5 Max Preview leads

Mixtral 8x22B: 19.9 (#248), Qwen3.5 Max Preview: 30.8 (#84)

Reasoning benchmarks
BenchmarkMixtral 8x22BQwen3.5 Max Preview
LMArena Hard Prompts11501483
DTBench55.1%—
Epoch Capabilities Index122.03—
ForecastBench56.3—

Math Qwen3.5 Max Preview leads

Mixtral 8x22B: 22.9 (#275), Qwen3.5 Max Preview: 40.1 (#94)

Math benchmarks
BenchmarkMixtral 8x22BQwen3.5 Max Preview
LMArena Math11841474
Omni-MATH16.3%—
MATH Level 524.2%—

Knowledge Qwen3.5 Max Preview leads

Mixtral 8x22B: 15.1 (#293), Qwen3.5 Max Preview: 41.8 (#107)

Knowledge benchmarks
BenchmarkMixtral 8x22BQwen3.5 Max Preview
LMArena Expert11131489
GPQA Diamond34.1%—
MMLU-Pro46%—
GPQA (HELM)33.4%—
MMLU77.8%—

Multilingual Qwen3.5 Max Preview leads

Mixtral 8x22B: 32.8 (#255), Qwen3.5 Max Preview: 56.2 (#22)

Multilingual benchmarks
BenchmarkMixtral 8x22BQwen3.5 Max Preview
LMArena Non-English11281465
LMArena Chinese11161534
LMArena French11661484
LMArena German11411487
LMArena Japanese10371495
LMArena Korean10571438
LMArena Russian11581471
LMArena Spanish11511470

Instruction Following Qwen3.5 Max Preview leads

Mixtral 8x22B: 57.7 (#266), Qwen3.5 Max Preview: 77.0 (#31)

Instruction Following benchmarks
BenchmarkMixtral 8x22BQwen3.5 Max Preview
LMArena Instruction Following11471467
IFEval72.4%—

Long Context Qwen3.5 Max Preview leads

Mixtral 8x22B: 34.7 (#247), Qwen3.5 Max Preview: 45.2 (#45)

Long Context benchmarks
BenchmarkMixtral 8x22BQwen3.5 Max Preview
LMArena Longer Query11441476

Writing & Preference Qwen3.5 Max Preview leads

Mixtral 8x22B: 36.9 (#262), Qwen3.5 Max Preview: 66.0 (#41)

Writing & Preference benchmarks
BenchmarkMixtral 8x22BQwen3.5 Max Preview
LMArena Text11621470
LMArena Creative Writing11411464
LMArena Multi-Turn11301478
WildBench71.1%—

Frequently asked questions

Is Mixtral 8x22B better than Qwen3.5 Max Preview?

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 27.1 on the Noometry Index.

Is Mixtral 8x22B or Qwen3.5 Max Preview better for coding?

Qwen3.5 Max Preview scores higher on coding benchmarks: 44.0 versus 24.2 in the Noometry coding category.

How many benchmarks do Mixtral 8x22B and Qwen3.5 Max Preview share?

17 benchmarks have published results for both models. Mixtral 8x22B has 34 scored results on Noometry and Qwen3.5 Max Preview has 17.

Related comparisons

Go deeper