Model comparison

Mistral vs Qwen1.5 4b Chat

Mistral is the stronger model overall, scoring 29.9 to 28.8 on the Noometry Index.

Last verified . 13 shared benchmarks.

Mistral Mistral AI

29.9

Rank #303 Confirmed

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Mistral scores higher in 6 categories and Qwen1.5 4b Chat in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral leads 37.0 to 23.8.
  • Qwen1.5 4b Chat has downloadable open weights; the other is API-only.

Side by side

Mistral and Qwen1.5 4b Chat specifications
MistralQwen1.5 4b Chat
ProviderMistral AIAlibaba (Qwen)
Noometry Index29.928.8
Released——
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2213

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral leads

Mistral: 33.8 (#250), Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkMistralQwen1.5 4b Chat
LMArena Coding1162999

Reasoning Mistral leads

Mistral: 22.2 (#200), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkMistralQwen1.5 4b Chat
LMArena Hard Prompts1149976

Math Qwen1.5 4b Chat leads

Mistral: 22.3 (#278), Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkMistralQwen1.5 4b Chat
LMArena Math11801026
Omni-MATH7.2%—

Knowledge Qwen1.5 4b Chat leads

Mistral: 16.6 (#288), Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkMistralQwen1.5 4b Chat
LMArena Expert1125980
MMLU-Pro27.7%—
GPQA (HELM)30.3%—

Multilingual Mistral leads

Mistral: 32.8 (#254), Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkMistralQwen1.5 4b Chat
LMArena Non-English1129979
LMArena Chinese11091024
LMArena German1155902
LMArena Russian1168952
LMArena French1180—
LMArena Japanese1013—
LMArena Korean1032—
LMArena Spanish1143—

Instruction Following Mistral leads

Mistral: 52.6 (#288), Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkMistralQwen1.5 4b Chat
LMArena Instruction Following1152978
IFEval56.8%—

Long Context Mistral leads

Mistral: 35.0 (#245), Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkMistralQwen1.5 4b Chat
LMArena Longer Query1153988

Writing & Preference Mistral leads

Mistral: 37.0 (#260), Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkMistralQwen1.5 4b Chat
LMArena Text1165997
LMArena Creative Writing1158969
LMArena Multi-Turn1147977
WildBench66%—

Frequently asked questions

Is Mistral better than Qwen1.5 4b Chat?

Mistral is the stronger model overall, scoring 29.9 to 28.8 on the Noometry Index.

Is Mistral or Qwen1.5 4b Chat better for coding?

Mistral scores higher on coding benchmarks: 33.8 versus 29.1 in the Noometry coding category.

How many benchmarks do Mistral and Qwen1.5 4b Chat share?

13 benchmarks have published results for both models. Mistral has 22 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper