Model comparison
Mistral Medium 3.1 vs Qwen3.5 Plus
Qwen3.5 Plus is the stronger model overall, scoring 42.9 to 31.9 on the Noometry Index.
Last verified . 0 shared benchmarks.
Summary
- The widest gap is in reasoning, where Qwen3.5 Plus leads 32.8 to 10.6.
- Mistral Medium 3.1 is cheaper at $0.40 / $2 per million input/output tokens, against $0.40 / $2.40 for Qwen3.5 Plus.
- Qwen3.5 Plus accepts more context: 1M tokens versus 131K.
Side by side
| Mistral Medium 3.1 | Qwen3.5 Plus | |
|---|---|---|
| Provider | Mistral AI | Alibaba (Qwen) |
| Noometry Index | 31.9 | 42.9 |
| Released | — | 2026-02-16 |
| Weights | Proprietary | Proprietary |
| Context window | 131K | 1M |
| Max output | 105K | 66K |
| Input $ / M tokens | $0.40 | $0.40 |
| Output $ / M tokens | $2 | $2.40 |
| Results tracked | 3 | 15 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Not comparable
Mistral Medium 3.1: —, Qwen3.5 Plus: —
| Benchmark | Mistral Medium 3.1 | Qwen3.5 Plus |
|---|---|---|
| ALE-Bench | — | 621.92 |
Agentic & Tool Use Not comparable
Mistral Medium 3.1: —, Qwen3.5 Plus: —
| Benchmark | Mistral Medium 3.1 | Qwen3.5 Plus |
|---|---|---|
| Vending-Bench 2 | — | 0.54 |
Reasoning Qwen3.5 Plus leads
Mistral Medium 3.1: 10.6 (#341), Qwen3.5 Plus: 32.8 (#74)
| Benchmark | Mistral Medium 3.1 | Qwen3.5 Plus |
|---|---|---|
| NYT Connections (extended) | 6.5% | — |
| Chess Puzzles | — | 22% |
| Thematic Generalization | 20.3% | — |
| Mystery Game Puzzles | — | 17% |
| DTBench | — | 80.5% |
| LMCA | — | 36.4% |
| Epoch Capabilities Index | — | 146.78 |
Math Not comparable
Mistral Medium 3.1: —, Qwen3.5 Plus: 49.6 (#61)
| Benchmark | Mistral Medium 3.1 | Qwen3.5 Plus |
|---|---|---|
| OTIS Mock AIME 2024-2025 | — | 86.7% |
| FrontierMath (Feb 2025 set) | — | 21% |
| FrontierMath Tier 4 (v1) | — | 2.1% |
Knowledge Not comparable
Mistral Medium 3.1: —, Qwen3.5 Plus: 46.0 (#83)
| Benchmark | Mistral Medium 3.1 | Qwen3.5 Plus |
|---|---|---|
| GPQA Diamond | — | 84.8% |
| SimpleQA Verified | — | 25.4% |
| Vectara Hallucination Rate | — | 10.7% |
Long Context Not comparable
Mistral Medium 3.1: —, Qwen3.5 Plus: 43.0 (#113)
| Benchmark | Mistral Medium 3.1 | Qwen3.5 Plus |
|---|---|---|
| CL-bench | — | 19.8% |
| CL-bench Life | — | 12.4% |
Writing & Preference Not comparable
Mistral Medium 3.1: 55.5 (#145), Qwen3.5 Plus: —
| Benchmark | Mistral Medium 3.1 | Qwen3.5 Plus |
|---|---|---|
| EQ-Bench Creative Writing | 1476 | — |
Frequently asked questions
Is Mistral Medium 3.1 better than Qwen3.5 Plus?
Qwen3.5 Plus is the stronger model overall, scoring 42.9 to 31.9 on the Noometry Index.
Which is cheaper, Mistral Medium 3.1 or Qwen3.5 Plus?
Mistral Medium 3.1 is cheaper. It lists at $0.40 per million input tokens and $2 per million output tokens; Qwen3.5 Plus lists at $0.40 and $2.40.
Which has the bigger context window?
Qwen3.5 Plus does, with 1M tokens against 131K.
How many benchmarks do Mistral Medium 3.1 and Qwen3.5 Plus share?
0 benchmarks have published results for both models. Mistral Medium 3.1 has 3 scored results on Noometry and Qwen3.5 Plus has 15.