Model comparison

Mistral vs Qwen Max

Qwen Max is the stronger model overall, scoring 34.7 to 29.9 on the Noometry Index.

Last verified . 17 shared benchmarks.

Mistral Mistral AI

29.9

Rank #303 Confirmed

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Mistral scores higher in 1 category and Qwen Max in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Qwen Max leads 66.5 to 52.6.

Side by side

Mistral and Qwen Max specifications
MistralQwen Max
ProviderMistral AIAlibaba (Qwen)
Noometry Index29.934.7
Released—2024-04-03
WeightsProprietaryProprietary
Context window—33K
Max output—8K
Input $ / M tokens—$1.60
Output $ / M tokens—$6.40
Results tracked2223

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral leads

Mistral: 33.8 (#250), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkMistralQwen Max
LMArena Coding11621288
Aider Polyglot—21.8%

Reasoning Qwen Max leads

Mistral: 22.2 (#200), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkMistralQwen Max
LMArena Hard Prompts11491269

Math Too close to call

Mistral: 22.3 (#278), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkMistralQwen Max
LMArena Math11801275
OTIS Mock AIME 2024-2025—16.1%
Omni-MATH7.2%—
MATH Level 5—67.2%
FrontierMath (Feb 2025 set)—1%

Knowledge Qwen Max leads

Mistral: 16.6 (#288), Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkMistralQwen Max
LMArena Expert11251248
GPQA Diamond—56.1%
MMLU-Pro27.7%—
GPQA (HELM)30.3%—

Multilingual Qwen Max leads

Mistral: 32.8 (#254), Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkMistralQwen Max
LMArena Non-English11291263
LMArena Chinese11091254
LMArena French11801330
LMArena German11551254
LMArena Japanese10131205
LMArena Korean10321142
LMArena Russian11681274
LMArena Spanish11431290

Instruction Following Qwen Max leads

Mistral: 52.6 (#288), Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkMistralQwen Max
LMArena Instruction Following11521262
IFEval56.8%—

Long Context Qwen Max leads

Mistral: 35.0 (#245), Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkMistralQwen Max
LMArena Longer Query11531288
Fiction.LiveBench—66.7%

Writing & Preference Qwen Max leads

Mistral: 37.0 (#260), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkMistralQwen Max
LMArena Text11651282
LMArena Creative Writing11581248
LMArena Multi-Turn11471277
WildBench66%—

Frequently asked questions

Is Mistral better than Qwen Max?

Qwen Max is the stronger model overall, scoring 34.7 to 29.9 on the Noometry Index.

Is Mistral or Qwen Max better for coding?

Mistral scores higher on coding benchmarks: 33.8 versus 30.7 in the Noometry coding category.

How many benchmarks do Mistral and Qwen Max share?

17 benchmarks have published results for both models. Mistral has 22 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper