Model comparison

Pixtral Large vs Qwen Max

Qwen Max is the stronger model overall, scoring 34.7 to 32.2 on the Noometry Index.

Last verified . 0 shared benchmarks.

Pixtral Large Mistral AI

32.2

Rank #259 Reported

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • The widest gap is in writing & preference, where Qwen Max leads 47.8 to 32.9.
  • Qwen Max is cheaper at $1.60 / $6.40 per million input/output tokens, against $2 / $6 for Pixtral Large.
  • Pixtral Large accepts more context: 128K tokens versus 33K.
  • Pixtral Large has downloadable open weights; the other is API-only.

Side by side

Pixtral Large and Qwen Max specifications
Pixtral LargeQwen Max
ProviderMistral AIAlibaba (Qwen)
Noometry Index32.234.7
Released2024-11-012024-04-03
WeightsOpenProprietary
Context window128K33K
Max output128K8K
Input $ / M tokens$2$1.60
Output $ / M tokens$6$6.40
Results tracked323

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Pixtral Large: —, Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkPixtral LargeQwen Max
Aider Polyglot—21.8%
LMArena Coding—1288

Reasoning Qwen Max leads

Pixtral Large: 21.7 (#218), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkPixtral LargeQwen Max
EnigmaEval0.8%—
LMArena Hard Prompts—1269

Math Not comparable

Pixtral Large: —, Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkPixtral LargeQwen Max
OTIS Mock AIME 2024-2025—16.1%
LMArena Math—1275
MATH Level 5—67.2%
FrontierMath (Feb 2025 set)—1%

Knowledge Not comparable

Pixtral Large: —, Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkPixtral LargeQwen Max
GPQA Diamond—56.1%
LMArena Expert—1248

Multimodal Not comparable

Pixtral Large: 30.6 (#111), Qwen Max: —

Multimodal benchmarks
BenchmarkPixtral LargeQwen Max
LMArena Vision1089—

Multilingual Not comparable

Pixtral Large: —, Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkPixtral LargeQwen Max
LMArena Non-English—1263
LMArena Chinese—1254
LMArena French—1330
LMArena German—1254
LMArena Japanese—1205
LMArena Korean—1142
LMArena Russian—1274
LMArena Spanish—1290

Instruction Following Not comparable

Pixtral Large: —, Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkPixtral LargeQwen Max
LMArena Instruction Following—1262

Long Context Not comparable

Pixtral Large: —, Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkPixtral LargeQwen Max
Fiction.LiveBench—66.7%
LMArena Longer Query—1288

Writing & Preference Qwen Max leads

Pixtral Large: 32.9 (#278), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkPixtral LargeQwen Max
LMArena Text—1282
LMArena Creative Writing—1248
EQ-Bench Creative Writing988—
LMArena Multi-Turn—1277

Frequently asked questions

Is Pixtral Large better than Qwen Max?

Qwen Max is the stronger model overall, scoring 34.7 to 32.2 on the Noometry Index.

Which is cheaper, Pixtral Large or Qwen Max?

Qwen Max is cheaper. It lists at $1.60 per million input tokens and $6.40 per million output tokens; Pixtral Large lists at $2 and $6.

Which has the bigger context window?

Pixtral Large does, with 128K tokens against 33K.

How many benchmarks do Pixtral Large and Qwen Max share?

0 benchmarks have published results for both models. Pixtral Large has 3 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper