Model comparison
Mistral vs Qwen Max
Qwen Max is the stronger model overall, scoring 34.7 to 29.9 on the Noometry Index.
Last verified . 17 shared benchmarks.
Summary
- They share 17 benchmarks with published results for both. Mistral scores higher in 1 category and Qwen Max in 7 categories; 7 gaps are clear of the uncertainty.
- The widest gap is in instruction following, where Qwen Max leads 66.5 to 52.6.
Side by side
| Mistral | Qwen Max | |
|---|---|---|
| Provider | Mistral AI | Alibaba (Qwen) |
| Noometry Index | 29.9 | 34.7 |
| Released | — | 2024-04-03 |
| Weights | Proprietary | Proprietary |
| Context window | — | 33K |
| Max output | — | 8K |
| Input $ / M tokens | — | $1.60 |
| Output $ / M tokens | — | $6.40 |
| Results tracked | 22 | 23 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Mistral leads
Mistral: 33.8 (#250), Qwen Max: 30.7 (#292)
| Benchmark | Mistral | Qwen Max |
|---|---|---|
| LMArena Coding | 1162 | 1288 |
| Aider Polyglot | — | 21.8% |
Reasoning Qwen Max leads
Mistral: 22.2 (#200), Qwen Max: 25.1 (#151)
| Benchmark | Mistral | Qwen Max |
|---|---|---|
| LMArena Hard Prompts | 1149 | 1269 |
Math Too close to call
Mistral: 22.3 (#278), Qwen Max: 22.3 (#276)
| Benchmark | Mistral | Qwen Max |
|---|---|---|
| LMArena Math | 1180 | 1275 |
| OTIS Mock AIME 2024-2025 | — | 16.1% |
| Omni-MATH | 7.2% | — |
| MATH Level 5 | — | 67.2% |
| FrontierMath (Feb 2025 set) | — | 1% |
Knowledge Qwen Max leads
Mistral: 16.6 (#288), Qwen Max: 30.3 (#228)
| Benchmark | Mistral | Qwen Max |
|---|---|---|
| LMArena Expert | 1125 | 1248 |
| GPQA Diamond | — | 56.1% |
| MMLU-Pro | 27.7% | — |
| GPQA (HELM) | 30.3% | — |
Multilingual Qwen Max leads
Mistral: 32.8 (#254), Qwen Max: 41.8 (#202)
| Benchmark | Mistral | Qwen Max |
|---|---|---|
| LMArena Non-English | 1129 | 1263 |
| LMArena Chinese | 1109 | 1254 |
| LMArena French | 1180 | 1330 |
| LMArena German | 1155 | 1254 |
| LMArena Japanese | 1013 | 1205 |
| LMArena Korean | 1032 | 1142 |
| LMArena Russian | 1168 | 1274 |
| LMArena Spanish | 1143 | 1290 |
Instruction Following Qwen Max leads
Mistral: 52.6 (#288), Qwen Max: 66.5 (#208)
| Benchmark | Mistral | Qwen Max |
|---|---|---|
| LMArena Instruction Following | 1152 | 1262 |
| IFEval | 56.8% | — |
Long Context Qwen Max leads
Mistral: 35.0 (#245), Qwen Max: 39.4 (#180)
| Benchmark | Mistral | Qwen Max |
|---|---|---|
| LMArena Longer Query | 1153 | 1288 |
| Fiction.LiveBench | — | 66.7% |
Writing & Preference Qwen Max leads
Mistral: 37.0 (#260), Qwen Max: 47.8 (#205)
| Benchmark | Mistral | Qwen Max |
|---|---|---|
| LMArena Text | 1165 | 1282 |
| LMArena Creative Writing | 1158 | 1248 |
| LMArena Multi-Turn | 1147 | 1277 |
| WildBench | 66% | — |
Frequently asked questions
Is Mistral better than Qwen Max?
Qwen Max is the stronger model overall, scoring 34.7 to 29.9 on the Noometry Index.
Is Mistral or Qwen Max better for coding?
Mistral scores higher on coding benchmarks: 33.8 versus 30.7 in the Noometry coding category.
How many benchmarks do Mistral and Qwen Max share?
17 benchmarks have published results for both models. Mistral has 22 scored results on Noometry and Qwen Max has 23.