Model comparison

Mistral Medium 3.1 vs Qwen1.5 4b Chat

Mistral Medium 3.1 is the stronger model overall, scoring 31.9 to 28.8 on the Noometry Index.

Last verified . 0 shared benchmarks.

Mistral Medium 3.1 Mistral AI

31.9

Rank #266 Reported

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • The widest gap is in writing & preference, where Mistral Medium 3.1 leads 55.5 to 23.8.
  • Qwen1.5 4b Chat has downloadable open weights; the other is API-only.

Side by side

Mistral Medium 3.1 and Qwen1.5 4b Chat specifications
Mistral Medium 3.1Qwen1.5 4b Chat
ProviderMistral AIAlibaba (Qwen)
Noometry Index31.928.8
Released——
WeightsProprietaryOpen
Context window131K—
Max output105K—
Input $ / M tokens$0.40—
Output $ / M tokens$2—
Results tracked313

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Mistral Medium 3.1: —, Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkMistral Medium 3.1Qwen1.5 4b Chat
LMArena Coding—999

Reasoning Qwen1.5 4b Chat leads

Mistral Medium 3.1: 10.6 (#341), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkMistral Medium 3.1Qwen1.5 4b Chat
NYT Connections (extended)6.5%—
Thematic Generalization20.3%—
LMArena Hard Prompts—976

Math Not comparable

Mistral Medium 3.1: —, Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkMistral Medium 3.1Qwen1.5 4b Chat
LMArena Math—1026

Knowledge Not comparable

Mistral Medium 3.1: —, Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkMistral Medium 3.1Qwen1.5 4b Chat
LMArena Expert—980

Multilingual Not comparable

Mistral Medium 3.1: —, Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkMistral Medium 3.1Qwen1.5 4b Chat
LMArena Non-English—979
LMArena Chinese—1024
LMArena German—902
LMArena Russian—952

Instruction Following Not comparable

Mistral Medium 3.1: —, Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkMistral Medium 3.1Qwen1.5 4b Chat
LMArena Instruction Following—978

Long Context Not comparable

Mistral Medium 3.1: —, Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkMistral Medium 3.1Qwen1.5 4b Chat
LMArena Longer Query—988

Writing & Preference Mistral Medium 3.1 leads

Mistral Medium 3.1: 55.5 (#145), Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkMistral Medium 3.1Qwen1.5 4b Chat
LMArena Text—997
LMArena Creative Writing—969
EQ-Bench Creative Writing1476—
LMArena Multi-Turn—977

Frequently asked questions

Is Mistral Medium 3.1 better than Qwen1.5 4b Chat?

Mistral Medium 3.1 is the stronger model overall, scoring 31.9 to 28.8 on the Noometry Index.

How many benchmarks do Mistral Medium 3.1 and Qwen1.5 4b Chat share?

0 benchmarks have published results for both models. Mistral Medium 3.1 has 3 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper