Model comparison

Mixtral 8x7B vs Qwen Plus

Qwen Plus is the stronger model overall, scoring 37.1 to 27.1 on the Noometry Index.

Last verified . 16 shared benchmarks.

Mixtral 8x7B Mistral AI

27.1

Rank #334 Confirmed

Qwen Plus Alibaba (Qwen)

37.1

Rank #210 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Mixtral 8x7B scores higher in 0 categories and Qwen Plus in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen Plus leads 52.2 to 34.2.
  • The biggest single-benchmark swing is MATH Level 5: 10% for Mixtral 8x7B and 65.3% for Qwen Plus.
  • Qwen Plus is cheaper at $0.40 / $1.20 per million input/output tokens, against $0.70 / $0.70 for Mixtral 8x7B.
  • Qwen Plus accepts more context: 1M tokens versus 32K.
  • Mixtral 8x7B has downloadable open weights; the other is API-only.

Side by side

Mixtral 8x7B and Qwen Plus specifications
Mixtral 8x7BQwen Plus
ProviderMistral AIAlibaba (Qwen)
Noometry Index27.137.1
Released2023-12-112024-01-25
WeightsOpenProprietary
Context window32K1M
Max output32K33K
Input $ / M tokens$0.70$0.40
Output $ / M tokens$0.70$1.20
Results tracked3820

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen Plus leads

Mixtral 8x7B: 32.8 (#269), Qwen Plus: 38.9 (#167)

Coding benchmarks
BenchmarkMixtral 8x7BQwen Plus
LMArena Coding11261328
HumanEval+39.6%—
MBPP+49.7%—

Reasoning Qwen Plus leads

Mixtral 8x7B: 18.2 (#285), Qwen Plus: 28.4 (#107)

Reasoning benchmarks
BenchmarkMixtral 8x7BQwen Plus
LMArena Hard Prompts11151317
DTBench49.6%81.1%
Kagi LLM Benchmark—63.3%
LMCA—24%
Adversarial NLI55.2%—
Epoch Capabilities Index118.47—
ForecastBench56.3—
HellaSwag86.7%—
PIQA83.6%—
WinoGrande77.2%—

Math Qwen Plus leads

Mixtral 8x7B: 18.8 (#289), Qwen Plus: 23.3 (#271)

Math benchmarks
BenchmarkMixtral 8x7BQwen Plus
LMArena Math11471326
MATH Level 510%65.3%
OTIS Mock AIME 2024-2025—17.8%
Omni-MATH10.5%—
FrontierMath (Feb 2025 set)—1.7%
GSM8K74.4%—

Knowledge Qwen Plus leads

Mixtral 8x7B: 11.0 (#301), Qwen Plus: 27.4 (#251)

Knowledge benchmarks
BenchmarkMixtral 8x7BQwen Plus
GPQA Diamond30.6%48.1%
LMArena Expert10881328
MMLU-Pro33.5%—
GPQA (HELM)29.6%—
ARC (AI2) Challenge87.3%—
MMLU70.6%—
OpenBookQA85.8%—
TriviaQA82.2%—

Multilingual Qwen Plus leads

Mixtral 8x7B: 29.6 (#266), Qwen Plus: 45.1 (#175)

Multilingual benchmarks
BenchmarkMixtral 8x7BQwen Plus
LMArena Non-English10771310
LMArena Chinese10551347
LMArena Japanese9311251
LMArena Russian10901323
LMArena French1166—
LMArena German1114—
LMArena Korean968—
LMArena Spanish1111—

Instruction Following Qwen Plus leads

Mixtral 8x7B: 51.0 (#297), Qwen Plus: 68.8 (#181)

Instruction Following benchmarks
BenchmarkMixtral 8x7BQwen Plus
LMArena Instruction Following11091303
IFEval57.5%—

Long Context Qwen Plus leads

Mixtral 8x7B: 33.4 (#260), Qwen Plus: 40.3 (#158)

Long Context benchmarks
BenchmarkMixtral 8x7BQwen Plus
LMArena Longer Query11031324

Writing & Preference Qwen Plus leads

Mixtral 8x7B: 34.2 (#270), Qwen Plus: 52.2 (#176)

Writing & Preference benchmarks
BenchmarkMixtral 8x7BQwen Plus
LMArena Text11321326
LMArena Creative Writing11091293
LMArena Multi-Turn11151336
WildBench67.3%—

Frequently asked questions

Is Mixtral 8x7B better than Qwen Plus?

Qwen Plus is the stronger model overall, scoring 37.1 to 27.1 on the Noometry Index.

Which is cheaper, Mixtral 8x7B or Qwen Plus?

Qwen Plus is cheaper. It lists at $0.40 per million input tokens and $1.20 per million output tokens; Mixtral 8x7B lists at $0.70 and $0.70.

Is Mixtral 8x7B or Qwen Plus better for coding?

Qwen Plus scores higher on coding benchmarks: 38.9 versus 32.8 in the Noometry coding category.

Which has the bigger context window?

Qwen Plus does, with 1M tokens against 32K.

How many benchmarks do Mixtral 8x7B and Qwen Plus share?

16 benchmarks have published results for both models. Mixtral 8x7B has 38 scored results on Noometry and Qwen Plus has 20.

Related comparisons

Go deeper