Model comparison

Mistral Medium 3.1 vs Qwen3.5 Plus

Qwen3.5 Plus is the stronger model overall, scoring 42.9 to 31.9 on the Noometry Index.

Last verified . 0 shared benchmarks.

Mistral Medium 3.1 Mistral AI

31.9

Rank #266 Reported

Qwen3.5 Plus Alibaba (Qwen)

42.9

Rank #106 Confirmed

Summary

  • The widest gap is in reasoning, where Qwen3.5 Plus leads 32.8 to 10.6.
  • Mistral Medium 3.1 is cheaper at $0.40 / $2 per million input/output tokens, against $0.40 / $2.40 for Qwen3.5 Plus.
  • Qwen3.5 Plus accepts more context: 1M tokens versus 131K.

Side by side

Mistral Medium 3.1 and Qwen3.5 Plus specifications
Mistral Medium 3.1Qwen3.5 Plus
ProviderMistral AIAlibaba (Qwen)
Noometry Index31.942.9
Released—2026-02-16
WeightsProprietaryProprietary
Context window131K1M
Max output105K66K
Input $ / M tokens$0.40$0.40
Output $ / M tokens$2$2.40
Results tracked315

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Mistral Medium 3.1: —, Qwen3.5 Plus: —

Coding benchmarks
BenchmarkMistral Medium 3.1Qwen3.5 Plus
ALE-Bench—621.92

Agentic & Tool Use Not comparable

Mistral Medium 3.1: —, Qwen3.5 Plus: —

Agentic & Tool Use benchmarks
BenchmarkMistral Medium 3.1Qwen3.5 Plus
Vending-Bench 2—0.54

Reasoning Qwen3.5 Plus leads

Mistral Medium 3.1: 10.6 (#341), Qwen3.5 Plus: 32.8 (#74)

Reasoning benchmarks
BenchmarkMistral Medium 3.1Qwen3.5 Plus
NYT Connections (extended)6.5%—
Chess Puzzles—22%
Thematic Generalization20.3%—
Mystery Game Puzzles—17%
DTBench—80.5%
LMCA—36.4%
Epoch Capabilities Index—146.78

Math Not comparable

Mistral Medium 3.1: —, Qwen3.5 Plus: 49.6 (#61)

Math benchmarks
BenchmarkMistral Medium 3.1Qwen3.5 Plus
OTIS Mock AIME 2024-2025—86.7%
FrontierMath (Feb 2025 set)—21%
FrontierMath Tier 4 (v1)—2.1%

Knowledge Not comparable

Mistral Medium 3.1: —, Qwen3.5 Plus: 46.0 (#83)

Knowledge benchmarks
BenchmarkMistral Medium 3.1Qwen3.5 Plus
GPQA Diamond—84.8%
SimpleQA Verified—25.4%
Vectara Hallucination Rate—10.7%

Long Context Not comparable

Mistral Medium 3.1: —, Qwen3.5 Plus: 43.0 (#113)

Long Context benchmarks
BenchmarkMistral Medium 3.1Qwen3.5 Plus
CL-bench—19.8%
CL-bench Life—12.4%

Writing & Preference Not comparable

Mistral Medium 3.1: 55.5 (#145), Qwen3.5 Plus: —

Writing & Preference benchmarks
BenchmarkMistral Medium 3.1Qwen3.5 Plus
EQ-Bench Creative Writing1476—

Frequently asked questions

Is Mistral Medium 3.1 better than Qwen3.5 Plus?

Qwen3.5 Plus is the stronger model overall, scoring 42.9 to 31.9 on the Noometry Index.

Which is cheaper, Mistral Medium 3.1 or Qwen3.5 Plus?

Mistral Medium 3.1 is cheaper. It lists at $0.40 per million input tokens and $2 per million output tokens; Qwen3.5 Plus lists at $0.40 and $2.40.

Which has the bigger context window?

Qwen3.5 Plus does, with 1M tokens against 131K.

How many benchmarks do Mistral Medium 3.1 and Qwen3.5 Plus share?

0 benchmarks have published results for both models. Mistral Medium 3.1 has 3 scored results on Noometry and Qwen3.5 Plus has 15.

Related comparisons

Go deeper