Model comparison

Mistral vs Qwen3.6 Plus

Qwen3.6 Plus is the stronger model overall, scoring 47.5 to 29.9 on the Noometry Index.

Last verified . 17 shared benchmarks.

Mistral Mistral AI

29.9

Rank #303 Confirmed

Qwen3.6 Plus Alibaba (Qwen)

47.5

Rank #62 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Mistral scores higher in 0 categories and Qwen3.6 Plus in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3.6 Plus leads 56.1 to 16.6.

Side by side

Mistral and Qwen3.6 Plus specifications
MistralQwen3.6 Plus
ProviderMistral AIAlibaba (Qwen)
Noometry Index29.947.5
Released—2026-03-31
WeightsProprietaryProprietary
Context window—1M
Max output—66K
Input $ / M tokens—$0.50
Output $ / M tokens—$3
Results tracked2237

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.6 Plus leads

Mistral: 33.8 (#250), Qwen3.6 Plus: 40.8 (#130)

Coding benchmarks
BenchmarkMistralQwen3.6 Plus
LMArena Coding11621467
SWE-bench Verified—57.9%
LMArena WebDev—1461
SciCode—40.7%
ALE-Bench—670.15

Agentic & Tool Use Not comparable

Mistral: —, Qwen3.6 Plus: —

Agentic & Tool Use benchmarks
BenchmarkMistralQwen3.6 Plus
Vending-Bench 2—5,115

Reasoning Qwen3.6 Plus leads

Mistral: 22.2 (#200), Qwen3.6 Plus: 29.3 (#93)

Reasoning benchmarks
BenchmarkMistralQwen3.6 Plus
LMArena Hard Prompts11491449
NYT Connections (extended)—60.3%
CritPt—2.9%
Chess Puzzles—17%
Thematic Generalization—59.5%
Mystery Game Puzzles—12%
DTBench—81.9%
LMCA—33.1%
Epoch Capabilities Index—147.65

Math Qwen3.6 Plus leads

Mistral: 22.3 (#278), Qwen3.6 Plus: 51.8 (#54)

Math benchmarks
BenchmarkMistralQwen3.6 Plus
LMArena Math11801450
FrontierMath (Tiers 1-3)—38.2%
OTIS Mock AIME 2024-2025—93.3%
Omni-MATH7.2%—
FrontierMath (Feb 2025 set)—26.2%
FrontierMath Tier 4 (v1)—8.3%

Knowledge Qwen3.6 Plus leads

Mistral: 16.6 (#288), Qwen3.6 Plus: 56.1 (#45)

Knowledge benchmarks
BenchmarkMistralQwen3.6 Plus
LMArena Expert11251454
GPQA Diamond—88.4%
SimpleQA Verified—44.1%
MMLU-Pro27.7%—
GPQA (HELM)30.3%—

Multilingual Qwen3.6 Plus leads

Mistral: 32.8 (#254), Qwen3.6 Plus: 53.3 (#70)

Multilingual benchmarks
BenchmarkMistralQwen3.6 Plus
LMArena Non-English11291424
LMArena Chinese11091477
LMArena French11801455
LMArena German11551452
LMArena Japanese10131389
LMArena Korean10321379
LMArena Russian11681434
LMArena Spanish11431432

Instruction Following Qwen3.6 Plus leads

Mistral: 52.6 (#288), Qwen3.6 Plus: 75.0 (#74)

Instruction Following benchmarks
BenchmarkMistralQwen3.6 Plus
LMArena Instruction Following11521425
IFEval56.8%—

Long Context Qwen3.6 Plus leads

Mistral: 35.0 (#245), Qwen3.6 Plus: 45.2 (#49)

Long Context benchmarks
BenchmarkMistralQwen3.6 Plus
LMArena Longer Query11531439
CL-bench—20.3%

Writing & Preference Qwen3.6 Plus leads

Mistral: 37.0 (#260), Qwen3.6 Plus: 62.2 (#82)

Writing & Preference benchmarks
BenchmarkMistralQwen3.6 Plus
LMArena Text11651437
LMArena Creative Writing11581404
LMArena Multi-Turn11471438
WildBench66%—

Frequently asked questions

Is Mistral better than Qwen3.6 Plus?

Qwen3.6 Plus is the stronger model overall, scoring 47.5 to 29.9 on the Noometry Index.

Is Mistral or Qwen3.6 Plus better for coding?

Qwen3.6 Plus scores higher on coding benchmarks: 40.8 versus 33.8 in the Noometry coding category.

How many benchmarks do Mistral and Qwen3.6 Plus share?

17 benchmarks have published results for both models. Mistral has 22 scored results on Noometry and Qwen3.6 Plus has 37.

Related comparisons

Go deeper