Model comparison
Mixtral 8x7B vs Qwen Plus
Qwen Plus is the stronger model overall, scoring 37.1 to 27.1 on the Noometry Index.
Last verified . 16 shared benchmarks.
Summary
- They share 16 benchmarks with published results for both. Mixtral 8x7B scores higher in 0 categories and Qwen Plus in 8 categories; 8 gaps are clear of the uncertainty.
- The widest gap is in writing & preference, where Qwen Plus leads 52.2 to 34.2.
- The biggest single-benchmark swing is MATH Level 5: 10% for Mixtral 8x7B and 65.3% for Qwen Plus.
- Qwen Plus is cheaper at $0.40 / $1.20 per million input/output tokens, against $0.70 / $0.70 for Mixtral 8x7B.
- Qwen Plus accepts more context: 1M tokens versus 32K.
- Mixtral 8x7B has downloadable open weights; the other is API-only.
Side by side
| Mixtral 8x7B | Qwen Plus | |
|---|---|---|
| Provider | Mistral AI | Alibaba (Qwen) |
| Noometry Index | 27.1 | 37.1 |
| Released | 2023-12-11 | 2024-01-25 |
| Weights | Open | Proprietary |
| Context window | 32K | 1M |
| Max output | 32K | 33K |
| Input $ / M tokens | $0.70 | $0.40 |
| Output $ / M tokens | $0.70 | $1.20 |
| Results tracked | 38 | 20 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Qwen Plus leads
Mixtral 8x7B: 32.8 (#269), Qwen Plus: 38.9 (#167)
| Benchmark | Mixtral 8x7B | Qwen Plus |
|---|---|---|
| LMArena Coding | 1126 | 1328 |
| HumanEval+ | 39.6% | — |
| MBPP+ | 49.7% | — |
Reasoning Qwen Plus leads
Mixtral 8x7B: 18.2 (#285), Qwen Plus: 28.4 (#107)
| Benchmark | Mixtral 8x7B | Qwen Plus |
|---|---|---|
| LMArena Hard Prompts | 1115 | 1317 |
| DTBench | 49.6% | 81.1% |
| Kagi LLM Benchmark | — | 63.3% |
| LMCA | — | 24% |
| Adversarial NLI | 55.2% | — |
| Epoch Capabilities Index | 118.47 | — |
| ForecastBench | 56.3 | — |
| HellaSwag | 86.7% | — |
| PIQA | 83.6% | — |
| WinoGrande | 77.2% | — |
Math Qwen Plus leads
Mixtral 8x7B: 18.8 (#289), Qwen Plus: 23.3 (#271)
| Benchmark | Mixtral 8x7B | Qwen Plus |
|---|---|---|
| LMArena Math | 1147 | 1326 |
| MATH Level 5 | 10% | 65.3% |
| OTIS Mock AIME 2024-2025 | — | 17.8% |
| Omni-MATH | 10.5% | — |
| FrontierMath (Feb 2025 set) | — | 1.7% |
| GSM8K | 74.4% | — |
Knowledge Qwen Plus leads
Mixtral 8x7B: 11.0 (#301), Qwen Plus: 27.4 (#251)
| Benchmark | Mixtral 8x7B | Qwen Plus |
|---|---|---|
| GPQA Diamond | 30.6% | 48.1% |
| LMArena Expert | 1088 | 1328 |
| MMLU-Pro | 33.5% | — |
| GPQA (HELM) | 29.6% | — |
| ARC (AI2) Challenge | 87.3% | — |
| MMLU | 70.6% | — |
| OpenBookQA | 85.8% | — |
| TriviaQA | 82.2% | — |
Multilingual Qwen Plus leads
Mixtral 8x7B: 29.6 (#266), Qwen Plus: 45.1 (#175)
| Benchmark | Mixtral 8x7B | Qwen Plus |
|---|---|---|
| LMArena Non-English | 1077 | 1310 |
| LMArena Chinese | 1055 | 1347 |
| LMArena Japanese | 931 | 1251 |
| LMArena Russian | 1090 | 1323 |
| LMArena French | 1166 | — |
| LMArena German | 1114 | — |
| LMArena Korean | 968 | — |
| LMArena Spanish | 1111 | — |
Instruction Following Qwen Plus leads
Mixtral 8x7B: 51.0 (#297), Qwen Plus: 68.8 (#181)
| Benchmark | Mixtral 8x7B | Qwen Plus |
|---|---|---|
| LMArena Instruction Following | 1109 | 1303 |
| IFEval | 57.5% | — |
Long Context Qwen Plus leads
Mixtral 8x7B: 33.4 (#260), Qwen Plus: 40.3 (#158)
| Benchmark | Mixtral 8x7B | Qwen Plus |
|---|---|---|
| LMArena Longer Query | 1103 | 1324 |
Writing & Preference Qwen Plus leads
Mixtral 8x7B: 34.2 (#270), Qwen Plus: 52.2 (#176)
| Benchmark | Mixtral 8x7B | Qwen Plus |
|---|---|---|
| LMArena Text | 1132 | 1326 |
| LMArena Creative Writing | 1109 | 1293 |
| LMArena Multi-Turn | 1115 | 1336 |
| WildBench | 67.3% | — |
Frequently asked questions
Is Mixtral 8x7B better than Qwen Plus?
Qwen Plus is the stronger model overall, scoring 37.1 to 27.1 on the Noometry Index.
Which is cheaper, Mixtral 8x7B or Qwen Plus?
Qwen Plus is cheaper. It lists at $0.40 per million input tokens and $1.20 per million output tokens; Mixtral 8x7B lists at $0.70 and $0.70.
Is Mixtral 8x7B or Qwen Plus better for coding?
Qwen Plus scores higher on coding benchmarks: 38.9 versus 32.8 in the Noometry coding category.
Which has the bigger context window?
Qwen Plus does, with 1M tokens against 32K.
How many benchmarks do Mixtral 8x7B and Qwen Plus share?
16 benchmarks have published results for both models. Mixtral 8x7B has 38 scored results on Noometry and Qwen Plus has 20.