Model comparison

Mistral vs Qwen3.5 Plus

Qwen3.5 Plus is the stronger model overall, scoring 42.9 to 29.9 on the Noometry Index.

Last verified . 0 shared benchmarks.

Mistral Mistral AI

29.9

Rank #303 Confirmed

Qwen3.5 Plus Alibaba (Qwen)

42.9

Rank #106 Confirmed

Summary

  • The widest gap is in knowledge, where Qwen3.5 Plus leads 46.0 to 16.6.

Side by side

Mistral and Qwen3.5 Plus specifications
MistralQwen3.5 Plus
ProviderMistral AIAlibaba (Qwen)
Noometry Index29.942.9
Released—2026-02-16
WeightsProprietaryProprietary
Context window—1M
Max output—66K
Input $ / M tokens—$0.40
Output $ / M tokens—$2.40
Results tracked2215

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Mistral: 33.8 (#250), Qwen3.5 Plus: —

Coding benchmarks
BenchmarkMistralQwen3.5 Plus
LMArena Coding1162—
ALE-Bench—621.92

Agentic & Tool Use Not comparable

Mistral: —, Qwen3.5 Plus: —

Agentic & Tool Use benchmarks
BenchmarkMistralQwen3.5 Plus
Vending-Bench 2—0.54

Reasoning Qwen3.5 Plus leads

Mistral: 22.2 (#200), Qwen3.5 Plus: 32.8 (#74)

Reasoning benchmarks
BenchmarkMistralQwen3.5 Plus
Chess Puzzles—22%
LMArena Hard Prompts1149—
Mystery Game Puzzles—17%
DTBench—80.5%
LMCA—36.4%
Epoch Capabilities Index—146.78

Math Qwen3.5 Plus leads

Mistral: 22.3 (#278), Qwen3.5 Plus: 49.6 (#61)

Math benchmarks
BenchmarkMistralQwen3.5 Plus
OTIS Mock AIME 2024-2025—86.7%
Omni-MATH7.2%—
LMArena Math1180—
FrontierMath (Feb 2025 set)—21%
FrontierMath Tier 4 (v1)—2.1%

Knowledge Qwen3.5 Plus leads

Mistral: 16.6 (#288), Qwen3.5 Plus: 46.0 (#83)

Knowledge benchmarks
BenchmarkMistralQwen3.5 Plus
GPQA Diamond—84.8%
SimpleQA Verified—25.4%
MMLU-Pro27.7%—
Vectara Hallucination Rate—10.7%
GPQA (HELM)30.3%—
LMArena Expert1125—

Multilingual Not comparable

Mistral: 32.8 (#254), Qwen3.5 Plus: —

Multilingual benchmarks
BenchmarkMistralQwen3.5 Plus
LMArena Non-English1129—
LMArena Chinese1109—
LMArena French1180—
LMArena German1155—
LMArena Japanese1013—
LMArena Korean1032—
LMArena Russian1168—
LMArena Spanish1143—

Instruction Following Not comparable

Mistral: 52.6 (#288), Qwen3.5 Plus: —

Instruction Following benchmarks
BenchmarkMistralQwen3.5 Plus
IFEval56.8%—
LMArena Instruction Following1152—

Long Context Qwen3.5 Plus leads

Mistral: 35.0 (#245), Qwen3.5 Plus: 43.0 (#113)

Long Context benchmarks
BenchmarkMistralQwen3.5 Plus
CL-bench—19.8%
CL-bench Life—12.4%
LMArena Longer Query1153—

Writing & Preference Not comparable

Mistral: 37.0 (#260), Qwen3.5 Plus: —

Writing & Preference benchmarks
BenchmarkMistralQwen3.5 Plus
LMArena Text1165—
LMArena Creative Writing1158—
WildBench66%—
LMArena Multi-Turn1147—

Frequently asked questions

Is Mistral better than Qwen3.5 Plus?

Qwen3.5 Plus is the stronger model overall, scoring 42.9 to 29.9 on the Noometry Index.

How many benchmarks do Mistral and Qwen3.5 Plus share?

0 benchmarks have published results for both models. Mistral has 22 scored results on Noometry and Qwen3.5 Plus has 15.

Related comparisons

Go deeper