Model comparison

Magistral Small vs Qwen3.5 397B-A17B

Qwen3.5 397B-A17B is the stronger model overall, scoring 46.0 to 30.2 on the Noometry Index. Magistral Small costs 1.8× less per token, which makes it the better buy when Qwen3.5 397B-A17B's lead doesn't matter for your workload.

Last verified . 6 shared benchmarks.

Magistral Small Mistral AI

30.2

Rank #296 Confirmed

Qwen3.5 397B-A17B Alibaba (Qwen)

46.0

Rank #67 Confirmed

Summary

  • They share 6 benchmarks with published results for both. Magistral Small scores higher in 0 categories and Qwen3.5 397B-A17B in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen3.5 397B-A17B leads 34.5 to 6.8.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 6.3% for Magistral Small and 73.7% for Qwen3.5 397B-A17B.
  • Magistral Small is cheaper at $0.50 / $1.50 per million input/output tokens, against $0.60 / $3.60 for Qwen3.5 397B-A17B.
  • Qwen3.5 397B-A17B accepts more context: 262K tokens versus 128K.

Side by side

Magistral Small and Qwen3.5 397B-A17B specifications
Magistral SmallQwen3.5 397B-A17B
ProviderMistral AIAlibaba (Qwen)
Noometry Index30.246.0
Released2025-06-102026-02-01
WeightsOpenOpen
Context window128K262K
Max output40K66K
Input $ / M tokens$0.50$0.60
Output $ / M tokens$1.50$3.60
Results tracked1036

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.5 397B-A17B leads

Magistral Small: 38.4 (#176), Qwen3.5 397B-A17B: 42.0 (#114)

Coding benchmarks
BenchmarkMagistral SmallQwen3.5 397B-A17B
LMArena WebDev—1400
SciCode35.2%—
LMArena Coding—1465

Agentic & Tool Use Not comparable

Magistral Small: —, Qwen3.5 397B-A17B: 33.3 (#53)

Agentic & Tool Use benchmarks
BenchmarkMagistral SmallQwen3.5 397B-A17B
APEX-Agents—24.9%
τ²-bench Airline—81.5%
τ²-bench Banking—9.8%
τ²-bench Retail—84.4%
τ²-bench Telecom—97.8%

Reasoning Qwen3.5 397B-A17B leads

Magistral Small: 6.8 (#350), Qwen3.5 397B-A17B: 34.5 (#70)

Reasoning benchmarks
BenchmarkMagistral SmallQwen3.5 397B-A17B
Kagi LLM Benchmark6.3%73.7%
Chess Puzzles3%13%
DTBench61.3%87.5%
Epoch Capabilities Index133.19146.65
ARC-AGI-20%—
NYT Connections (extended)—58.9%
ARC-AGI-15%—
CritPt0.3%—
Thematic Generalization—65.1%
LMArena Hard Prompts—1448
Mystery Game Puzzles—18%
LMCA—37.9%

Math Qwen3.5 397B-A17B leads

Magistral Small: 26.2 (#261), Qwen3.5 397B-A17B: 46.1 (#73)

Math benchmarks
BenchmarkMagistral SmallQwen3.5 397B-A17B
OTIS Mock AIME 2024-202530%88.9%
FrontierMath (Tiers 1-3)—31.2%
LMArena Math—1454

Knowledge Qwen3.5 397B-A17B leads

Magistral Small: 30.9 (#223), Qwen3.5 397B-A17B: 53.3 (#58)

Knowledge benchmarks
BenchmarkMagistral SmallQwen3.5 397B-A17B
GPQA Diamond56.1%86.4%
LMArena Expert—1462

Multimodal Not comparable

Magistral Small: —, Qwen3.5 397B-A17B: 40.7 (#44)

Multimodal benchmarks
BenchmarkMagistral SmallQwen3.5 397B-A17B
LMArena Vision—1263

Multilingual Not comparable

Magistral Small: —, Qwen3.5 397B-A17B: 53.7 (#59)

Multilingual benchmarks
BenchmarkMagistral SmallQwen3.5 397B-A17B
LMArena Non-English—1430
LMArena Chinese—1500
LMArena French—1461
LMArena German—1447
LMArena Japanese—1426
LMArena Korean—1384
LMArena Russian—1429
LMArena Spanish—1441

Instruction Following Not comparable

Magistral Small: —, Qwen3.5 397B-A17B: 75.0 (#77)

Instruction Following benchmarks
BenchmarkMagistral SmallQwen3.5 397B-A17B
LMArena Instruction Following—1424

Long Context Not comparable

Magistral Small: —, Qwen3.5 397B-A17B: 44.1 (#74)

Long Context benchmarks
BenchmarkMagistral SmallQwen3.5 397B-A17B
LMArena Longer Query—1442

Writing & Preference Not comparable

Magistral Small: —, Qwen3.5 397B-A17B: 62.3 (#79)

Writing & Preference benchmarks
BenchmarkMagistral SmallQwen3.5 397B-A17B
LMArena Text—1438
LMArena Creative Writing—1401
EQ-Bench Creative Writing—1478
LMArena Multi-Turn—1446

Frequently asked questions

Is Magistral Small better than Qwen3.5 397B-A17B?

Qwen3.5 397B-A17B is the stronger model overall, scoring 46.0 to 30.2 on the Noometry Index. Magistral Small costs 1.8× less per token, which makes it the better buy when Qwen3.5 397B-A17B's lead doesn't matter for your workload.

Which is cheaper, Magistral Small or Qwen3.5 397B-A17B?

Magistral Small is cheaper. It lists at $0.50 per million input tokens and $1.50 per million output tokens; Qwen3.5 397B-A17B lists at $0.60 and $3.60.

Is Magistral Small or Qwen3.5 397B-A17B better for coding?

Qwen3.5 397B-A17B scores higher on coding benchmarks: 42.0 versus 38.4 in the Noometry coding category.

Which has the bigger context window?

Qwen3.5 397B-A17B does, with 262K tokens against 128K.

How many benchmarks do Magistral Small and Qwen3.5 397B-A17B share?

6 benchmarks have published results for both models. Magistral Small has 10 scored results on Noometry and Qwen3.5 397B-A17B has 36.

Related comparisons

Go deeper