Model comparison

Magistral Medium vs Qwen3-4B

Magistral Medium is the stronger model overall, scoring 35.2 to 31.9 on the Noometry Index.

Last verified . 0 shared benchmarks.

Magistral Medium Mistral AI

35.2

Rank #227 Confirmed

Qwen3-4B Alibaba (Qwen)

31.9

Rank #264 Confirmed

Summary

  • The widest gap is in reasoning, where Qwen3-4B leads 19.2 to 8.6.

Side by side

Magistral Medium and Qwen3-4B specifications
Magistral MediumQwen3-4B
ProviderMistral AIAlibaba (Qwen)
Noometry Index35.231.9
Released2025-03-172025-04-29
WeightsOpenOpen
Context window262K—
Max output16K—
Input $ / M tokens$2—
Output $ / M tokens$5—
Results tracked226

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Magistral Medium: 39.1 (#161), Qwen3-4B: —

Coding benchmarks
BenchmarkMagistral MediumQwen3-4B
SciCode39.2%—
LMArena Coding1319—

Agentic & Tool Use Not comparable

Magistral Medium: —, Qwen3-4B: 27.6 (#100)

Agentic & Tool Use benchmarks
BenchmarkMagistral MediumQwen3-4B
Berkeley Function Calling Leaderboard—35.7%

Reasoning Qwen3-4B leads

Magistral Medium: 8.6 (#348), Qwen3-4B: 19.2 (#268)

Reasoning benchmarks
BenchmarkMagistral MediumQwen3-4B
ARC-AGI-20%—
Kagi LLM Benchmark16.2%—
ARC-AGI-16.1%—
CritPt0.3%—
Chess Puzzles—4%
LMArena Hard Prompts1267—

Math Magistral Medium leads

Magistral Medium: 35.1 (#189), Qwen3-4B: 29.7 (#240)

Math benchmarks
BenchmarkMagistral MediumQwen3-4B
MathArena Final-Answer Competitions—38.5%
OTIS Mock AIME 2024-2025—52.2%
LMArena Math1250—

Knowledge Too close to call

Magistral Medium: 33.5 (#202), Qwen3-4B: 33.0 (#208)

Knowledge benchmarks
BenchmarkMagistral MediumQwen3-4B
GPQA Diamond—52.3%
Vectara Hallucination Rate—5.7%
LMArena Expert1223—

Multilingual Not comparable

Magistral Medium: 39.6 (#224), Qwen3-4B: —

Multilingual benchmarks
BenchmarkMagistral MediumQwen3-4B
LMArena Non-English1232—
LMArena Chinese1227—
LMArena French1267—
LMArena German1248—
LMArena Japanese1175—
LMArena Korean1125—
LMArena Russian1224—
LMArena Spanish1271—

Instruction Following Not comparable

Magistral Medium: 66.0 (#211), Qwen3-4B: —

Instruction Following benchmarks
BenchmarkMagistral MediumQwen3-4B
LMArena Instruction Following1254—

Long Context Not comparable

Magistral Medium: 39.3 (#183), Qwen3-4B: —

Long Context benchmarks
BenchmarkMagistral MediumQwen3-4B
LMArena Longer Query1295—

Writing & Preference Not comparable

Magistral Medium: 46.3 (#219), Qwen3-4B: —

Writing & Preference benchmarks
BenchmarkMagistral MediumQwen3-4B
LMArena Text1255—
LMArena Creative Writing1245—
LMArena Multi-Turn1275—

Frequently asked questions

Is Magistral Medium better than Qwen3-4B?

Magistral Medium is the stronger model overall, scoring 35.2 to 31.9 on the Noometry Index.

How many benchmarks do Magistral Medium and Qwen3-4B share?

0 benchmarks have published results for both models. Magistral Medium has 22 scored results on Noometry and Qwen3-4B has 6.

Related comparisons

Go deeper