Model comparison

Mistral Large vs Qwen3.5 Max Preview

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 31.9 on the Noometry Index.

Last verified . 17 shared benchmarks.

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Qwen3.5 Max Preview Alibaba (Qwen)

45.3

Rank #71 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Mistral Large scores higher in 0 categories and Qwen3.5 Max Preview in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen3.5 Max Preview leads 66.0 to 40.7.
  • Mistral Large has downloadable open weights; the other is API-only.

Side by side

Mistral Large and Qwen3.5 Max Preview specifications
Mistral LargeQwen3.5 Max Preview
ProviderMistral AIAlibaba (Qwen)
Noometry Index31.945.3
Released2024-02-26—
WeightsOpenProprietary
Context window131K—
Max output16K—
Input $ / M tokens$2—
Output $ / M tokens$6—
Results tracked5117

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.5 Max Preview leads

Mistral Large: 34.3 (#240), Qwen3.5 Max Preview: 44.0 (#77)

Coding benchmarks
BenchmarkMistral LargeQwen3.5 Max Preview
LMArena Coding12771487
SciCode36.2%—
BigCodeBench Instruct30%—
LiveBench Coding47.1%—
BigCodeBench Complete38.3%—
ALE-Bench264.7—
HumanEval+62.2%—
MBPP+59.5%—

Agentic & Tool Use Not comparable

Mistral Large: 28.6 (#89), Qwen3.5 Max Preview: —

Agentic & Tool Use benchmarks
BenchmarkMistral LargeQwen3.5 Max Preview
Berkeley Function Calling Leaderboard38.4%—

Reasoning Qwen3.5 Max Preview leads

Mistral Large: 15.8 (#310), Qwen3.5 Max Preview: 30.8 (#84)

Reasoning benchmarks
BenchmarkMistral LargeQwen3.5 Max Preview
LMArena Hard Prompts12571483
SimpleBench22.5%—
CritPt0%—
LiveBench Reasoning43.5%—
DTBench65.1%—
LiveBench Data Analysis50.1%—
LMCA16.7%—
Epoch Capabilities Index128.52—
ForecastBench57.1—
LiveBench48.4%—

Math Qwen3.5 Max Preview leads

Mistral Large: 18.2 (#291), Qwen3.5 Max Preview: 40.1 (#94)

Math benchmarks
BenchmarkMistral LargeQwen3.5 Max Preview
LMArena Math12621474
OTIS Mock AIME 2024-20258.5%—
Omni-MATH28.1%—
LiveBench Math42.5%—
MATH Level 550.3%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Qwen3.5 Max Preview leads

Mistral Large: 30.1 (#230), Qwen3.5 Max Preview: 41.8 (#107)

Knowledge benchmarks
BenchmarkMistral LargeQwen3.5 Max Preview
LMArena Expert12321489
GPQA Diamond51.3%—
MMLU-Pro59.9%—
Confabulations21.4%—
Vectara Hallucination Rate4.5%—
GPQA (HELM)43.5%—
MMLU80%—

Multilingual Qwen3.5 Max Preview leads

Mistral Large: 40.0 (#219), Qwen3.5 Max Preview: 56.2 (#22)

Multilingual benchmarks
BenchmarkMistral LargeQwen3.5 Max Preview
LMArena Non-English12371465
LMArena Chinese12401534
LMArena French13251484
LMArena German12541487
LMArena Japanese11881495
LMArena Korean12021438
LMArena Russian12571471
LMArena Spanish12681470

Instruction Following Qwen3.5 Max Preview leads

Mistral Large: 67.9 (#191), Qwen3.5 Max Preview: 77.0 (#31)

Instruction Following benchmarks
BenchmarkMistral LargeQwen3.5 Max Preview
LMArena Instruction Following12491467
LiveBench Instruction Following67.9%—
IFEval87.7%—

Long Context Qwen3.5 Max Preview leads

Mistral Large: 38.3 (#199), Qwen3.5 Max Preview: 45.2 (#45)

Long Context benchmarks
BenchmarkMistral LargeQwen3.5 Max Preview
LMArena Longer Query12611476

Writing & Preference Qwen3.5 Max Preview leads

Mistral Large: 40.7 (#242), Qwen3.5 Max Preview: 66.0 (#41)

Writing & Preference benchmarks
BenchmarkMistral LargeQwen3.5 Max Preview
LMArena Text12661470
LMArena Creative Writing12431464
LMArena Multi-Turn12601478
Short-Story Creative Writing69%—
EQ-Bench Creative Writing985—
WildBench80.1%—
LiveBench Language39.4%—

Frequently asked questions

Is Mistral Large better than Qwen3.5 Max Preview?

Qwen3.5 Max Preview is the stronger model overall, scoring 45.3 to 31.9 on the Noometry Index.

Is Mistral Large or Qwen3.5 Max Preview better for coding?

Qwen3.5 Max Preview scores higher on coding benchmarks: 44.0 versus 34.3 in the Noometry coding category.

How many benchmarks do Mistral Large and Qwen3.5 Max Preview share?

17 benchmarks have published results for both models. Mistral Large has 51 scored results on Noometry and Qwen3.5 Max Preview has 17.

Related comparisons

Go deeper