Model comparison

Magistral Small vs Qwen1.5 4b Chat

Magistral Small is the stronger model overall, scoring 30.2 to 28.8 on the Noometry Index.

Last verified . 0 shared benchmarks.

Magistral Small Mistral AI

30.2

Rank #296 Confirmed

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • The widest gap is in reasoning, where Qwen1.5 4b Chat leads 18.5 to 6.8.

Side by side

Magistral Small and Qwen1.5 4b Chat specifications
Magistral SmallQwen1.5 4b Chat
ProviderMistral AIAlibaba (Qwen)
Noometry Index30.228.8
Released2025-06-10—
WeightsOpenOpen
Context window128K—
Max output40K—
Input $ / M tokens$0.50—
Output $ / M tokens$1.50—
Results tracked1013

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Small leads

Magistral Small: 38.4 (#176), Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkMagistral SmallQwen1.5 4b Chat
SciCode35.2%—
LMArena Coding—999

Reasoning Qwen1.5 4b Chat leads

Magistral Small: 6.8 (#350), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkMagistral SmallQwen1.5 4b Chat
ARC-AGI-20%—
Kagi LLM Benchmark6.3%—
ARC-AGI-15%—
CritPt0.3%—
Chess Puzzles3%—
LMArena Hard Prompts—976
DTBench61.3%—
Epoch Capabilities Index133.19—

Math Qwen1.5 4b Chat leads

Magistral Small: 26.2 (#261), Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkMagistral SmallQwen1.5 4b Chat
OTIS Mock AIME 2024-202530%—
LMArena Math—1026

Knowledge Magistral Small leads

Magistral Small: 30.9 (#223), Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkMagistral SmallQwen1.5 4b Chat
GPQA Diamond56.1%—
LMArena Expert—980

Multilingual Not comparable

Magistral Small: —, Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkMagistral SmallQwen1.5 4b Chat
LMArena Non-English—979
LMArena Chinese—1024
LMArena German—902
LMArena Russian—952

Instruction Following Not comparable

Magistral Small: —, Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkMagistral SmallQwen1.5 4b Chat
LMArena Instruction Following—978

Long Context Not comparable

Magistral Small: —, Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkMagistral SmallQwen1.5 4b Chat
LMArena Longer Query—988

Writing & Preference Not comparable

Magistral Small: —, Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkMagistral SmallQwen1.5 4b Chat
LMArena Text—997
LMArena Creative Writing—969
LMArena Multi-Turn—977

Frequently asked questions

Is Magistral Small better than Qwen1.5 4b Chat?

Magistral Small is the stronger model overall, scoring 30.2 to 28.8 on the Noometry Index.

Is Magistral Small or Qwen1.5 4b Chat better for coding?

Magistral Small scores higher on coding benchmarks: 38.4 versus 29.1 in the Noometry coding category.

How many benchmarks do Magistral Small and Qwen1.5 4b Chat share?

0 benchmarks have published results for both models. Magistral Small has 10 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper