Model comparison

Mixtral 8x22B vs Qwen Max

Qwen Max is the stronger model overall, scoring 34.7 to 27.1 on the Noometry Index.

Last verified . 19 shared benchmarks.

Mixtral 8x22B Mistral AI

27.1

Rank #333 Confirmed

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Mixtral 8x22B scores higher in 1 category and Qwen Max in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen Max leads 30.3 to 15.1.
  • The biggest single-benchmark swing is MATH Level 5: 24.2% for Mixtral 8x22B and 67.2% for Qwen Max.
  • Qwen Max is cheaper at $1.60 / $6.40 per million input/output tokens, against $2 / $6 for Mixtral 8x22B.
  • Mixtral 8x22B accepts more context: 64K tokens versus 33K.
  • Mixtral 8x22B has downloadable open weights; the other is API-only.

Side by side

Mixtral 8x22B and Qwen Max specifications
Mixtral 8x22BQwen Max
ProviderMistral AIAlibaba (Qwen)
Noometry Index27.134.7
Released2024-04-172024-04-03
WeightsOpenProprietary
Context window64K33K
Max output64K8K
Input $ / M tokens$2$1.60
Output $ / M tokens$6$6.40
Results tracked3423

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen Max leads

Mixtral 8x22B: 24.2 (#329), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkMixtral 8x22BQwen Max
LMArena Coding11661288
Aider Polyglot—21.8%
WeirdML3.2%—
BigCodeBench Instruct40.6%—
BigCodeBench Complete50.2%—
HumanEval+72%—
MBPP+64.3%—

Agentic & Tool Use Not comparable

Mixtral 8x22B: 23.1 (#127), Qwen Max: —

Agentic & Tool Use benchmarks
BenchmarkMixtral 8x22BQwen Max
Cybench7.5%—

Reasoning Qwen Max leads

Mixtral 8x22B: 19.9 (#248), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkMixtral 8x22BQwen Max
LMArena Hard Prompts11501269
DTBench55.1%—
Epoch Capabilities Index122.03—
ForecastBench56.3—

Math Too close to call

Mixtral 8x22B: 22.9 (#275), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkMixtral 8x22BQwen Max
LMArena Math11841275
MATH Level 524.2%67.2%
OTIS Mock AIME 2024-2025—16.1%
Omni-MATH16.3%—
FrontierMath (Feb 2025 set)—1%

Knowledge Qwen Max leads

Mixtral 8x22B: 15.1 (#293), Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkMixtral 8x22BQwen Max
GPQA Diamond34.1%56.1%
LMArena Expert11131248
MMLU-Pro46%—
GPQA (HELM)33.4%—
MMLU77.8%—

Multilingual Qwen Max leads

Mixtral 8x22B: 32.8 (#255), Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkMixtral 8x22BQwen Max
LMArena Non-English11281263
LMArena Chinese11161254
LMArena French11661330
LMArena German11411254
LMArena Japanese10371205
LMArena Korean10571142
LMArena Russian11581274
LMArena Spanish11511290

Instruction Following Qwen Max leads

Mixtral 8x22B: 57.7 (#266), Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkMixtral 8x22BQwen Max
LMArena Instruction Following11471262
IFEval72.4%—

Long Context Qwen Max leads

Mixtral 8x22B: 34.7 (#247), Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkMixtral 8x22BQwen Max
LMArena Longer Query11441288
Fiction.LiveBench—66.7%

Writing & Preference Qwen Max leads

Mixtral 8x22B: 36.9 (#262), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkMixtral 8x22BQwen Max
LMArena Text11621282
LMArena Creative Writing11411248
LMArena Multi-Turn11301277
WildBench71.1%—

Frequently asked questions

Is Mixtral 8x22B better than Qwen Max?

Qwen Max is the stronger model overall, scoring 34.7 to 27.1 on the Noometry Index.

Which is cheaper, Mixtral 8x22B or Qwen Max?

Qwen Max is cheaper. It lists at $1.60 per million input tokens and $6.40 per million output tokens; Mixtral 8x22B lists at $2 and $6.

Is Mixtral 8x22B or Qwen Max better for coding?

Qwen Max scores higher on coding benchmarks: 30.7 versus 24.2 in the Noometry coding category.

Which has the bigger context window?

Mixtral 8x22B does, with 64K tokens against 33K.

How many benchmarks do Mixtral 8x22B and Qwen Max share?

19 benchmarks have published results for both models. Mixtral 8x22B has 34 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper