Model comparison

Magistral Small vs Qwen3.6 Max Preview

Qwen3.6 Max Preview is the stronger model overall, scoring 51.5 to 30.2 on the Noometry Index. Magistral Small costs 3.9× less per token, which makes it the better buy when Qwen3.6 Max Preview's lead doesn't matter for your workload.

Last verified . 5 shared benchmarks.

Magistral Small Mistral AI

30.2

Rank #296 Confirmed

Qwen3.6 Max Preview Alibaba (Qwen)

51.5

Rank #43 Confirmed

Summary

  • They share 5 benchmarks with published results for both. Magistral Small scores higher in 0 categories and Qwen3.6 Max Preview in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen3.6 Max Preview leads 41.7 to 6.8.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 30% for Magistral Small and 91.1% for Qwen3.6 Max Preview.
  • Magistral Small is cheaper at $0.50 / $1.50 per million input/output tokens, against $1.30 / $7.80 for Qwen3.6 Max Preview.
  • Qwen3.6 Max Preview accepts more context: 262K tokens versus 128K.
  • Magistral Small has downloadable open weights; the other is API-only.

Side by side

Magistral Small and Qwen3.6 Max Preview specifications
Magistral SmallQwen3.6 Max Preview
ProviderMistral AIAlibaba (Qwen)
Noometry Index30.251.5
Released2025-06-102026-04-20
WeightsOpenProprietary
Context window128K262K
Max output40K66K
Input $ / M tokens$0.50$1.30
Output $ / M tokens$1.50$7.80
Results tracked1029

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.6 Max Preview leads

Magistral Small: 38.4 (#176), Qwen3.6 Max Preview: 48.7 (#54)

Coding benchmarks
BenchmarkMagistral SmallQwen3.6 Max Preview
SWE-bench Verified—76.7%
LMArena WebDev—1482
SciCode35.2%—
LMArena Coding—1471

Agentic & Tool Use Not comparable

Magistral Small: —, Qwen3.6 Max Preview: —

Agentic & Tool Use benchmarks
BenchmarkMagistral SmallQwen3.6 Max Preview
Vending-Bench 2—4,254

Reasoning Qwen3.6 Max Preview leads

Magistral Small: 6.8 (#350), Qwen3.6 Max Preview: 41.7 (#53)

Reasoning benchmarks
BenchmarkMagistral SmallQwen3.6 Max Preview
Chess Puzzles3%20%
DTBench61.3%87.2%
Epoch Capabilities Index133.19149.24
ARC-AGI-20%—
SimpleBench—63%
Kagi LLM Benchmark6.3%—
NYT Connections (extended)—74.1%
ARC-AGI-15%—
CritPt0.3%—
LMArena Hard Prompts—1457
Mystery Game Puzzles—19%
LMCA—42.5%

Math Qwen3.6 Max Preview leads

Magistral Small: 26.2 (#261), Qwen3.6 Max Preview: 54.1 (#46)

Math benchmarks
BenchmarkMagistral SmallQwen3.6 Max Preview
OTIS Mock AIME 2024-202530%91.1%
LMArena Math—1465
FrontierMath (Feb 2025 set)—23.1%
FrontierMath Tier 4 (v1)—4.2%

Knowledge Qwen3.6 Max Preview leads

Magistral Small: 30.9 (#223), Qwen3.6 Max Preview: 57.6 (#39)

Knowledge benchmarks
BenchmarkMagistral SmallQwen3.6 Max Preview
GPQA Diamond56.1%87.4%
SimpleQA Verified—52%
LMArena Expert—1478

Multilingual Not comparable

Magistral Small: —, Qwen3.6 Max Preview: 54.2 (#48)

Multilingual benchmarks
BenchmarkMagistral SmallQwen3.6 Max Preview
LMArena Non-English—1437
LMArena Chinese—1487
LMArena French—1449
LMArena Russian—1445
LMArena Spanish—1454

Instruction Following Not comparable

Magistral Small: —, Qwen3.6 Max Preview: 75.7 (#55)

Instruction Following benchmarks
BenchmarkMagistral SmallQwen3.6 Max Preview
LMArena Instruction Following—1438

Long Context Not comparable

Magistral Small: —, Qwen3.6 Max Preview: 44.6 (#61)

Long Context benchmarks
BenchmarkMagistral SmallQwen3.6 Max Preview
LMArena Longer Query—1457

Writing & Preference Not comparable

Magistral Small: —, Qwen3.6 Max Preview: 63.8 (#60)

Writing & Preference benchmarks
BenchmarkMagistral SmallQwen3.6 Max Preview
LMArena Text—1447
LMArena Creative Writing—1435
LMArena Multi-Turn—1456

Frequently asked questions

Is Magistral Small better than Qwen3.6 Max Preview?

Qwen3.6 Max Preview is the stronger model overall, scoring 51.5 to 30.2 on the Noometry Index. Magistral Small costs 3.9× less per token, which makes it the better buy when Qwen3.6 Max Preview's lead doesn't matter for your workload.

Which is cheaper, Magistral Small or Qwen3.6 Max Preview?

Magistral Small is cheaper. It lists at $0.50 per million input tokens and $1.50 per million output tokens; Qwen3.6 Max Preview lists at $1.30 and $7.80.

Is Magistral Small or Qwen3.6 Max Preview better for coding?

Qwen3.6 Max Preview scores higher on coding benchmarks: 48.7 versus 38.4 in the Noometry coding category.

Which has the bigger context window?

Qwen3.6 Max Preview does, with 262K tokens against 128K.

How many benchmarks do Magistral Small and Qwen3.6 Max Preview share?

5 benchmarks have published results for both models. Magistral Small has 10 scored results on Noometry and Qwen3.6 Max Preview has 29.

Related comparisons

Go deeper