Model comparison

Mistral Small vs Qwen1.5 4b Chat

Mistral Small is the stronger model overall, scoring 33.4 to 28.8 on the Noometry Index.

Last verified . 13 shared benchmarks.

Mistral Small Mistral AI

33.4

Rank #243 Confirmed

Qwen1.5 4b Chat Alibaba (Qwen)

28.8

Rank #322 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Mistral Small scores higher in 7 categories and Qwen1.5 4b Chat in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Small leads 52.5 to 23.8.

Side by side

Mistral Small and Qwen1.5 4b Chat specifications
Mistral SmallQwen1.5 4b Chat
ProviderMistral AIAlibaba (Qwen)
Noometry Index33.428.8
Released2024-02-26—
WeightsOpenOpen
Context window262K—
Max output256K—
Input $ / M tokens$0.15—
Output $ / M tokens$0.60—
Results tracked3913

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Small leads

Mistral Small: 34.0 (#247), Qwen1.5 4b Chat: 29.1 (#308)

Coding benchmarks
BenchmarkMistral SmallQwen1.5 4b Chat
LMArena Coding1362999
SciCode26.5%—
BigCodeBench Instruct36.1%—
LiveBench Coding36.2%—
BigCodeBench Complete46.6%—
ALE-Bench497.62—

Agentic & Tool Use Not comparable

Mistral Small: 28.1 (#93), Qwen1.5 4b Chat: —

Agentic & Tool Use benchmarks
BenchmarkMistral SmallQwen1.5 4b Chat
Berkeley Function Calling Leaderboard37.1%—

Reasoning Mistral Small leads

Mistral Small: 19.8 (#250), Qwen1.5 4b Chat: 18.5 (#279)

Reasoning benchmarks
BenchmarkMistral SmallQwen1.5 4b Chat
LMArena Hard Prompts1335976
Kagi LLM Benchmark37.8%—
CritPt0%—
LiveBench Reasoning44.8%—
DTBench70.9%—
LiveBench Data Analysis53.7%—
LMCA20.6%—
LiveBench44%—

Math Qwen1.5 4b Chat leads

Mistral Small: 16.4 (#293), Qwen1.5 4b Chat: 30.4 (#234)

Math benchmarks
BenchmarkMistral SmallQwen1.5 4b Chat
LMArena Math13411026
OTIS Mock AIME 2024-20255.8%—
LiveBench Math39.9%—
MATH Level 546.8%—

Knowledge Mistral Small leads

Mistral Small: 31.0 (#222), Qwen1.5 4b Chat: 26.7 (#255)

Knowledge benchmarks
BenchmarkMistral SmallQwen1.5 4b Chat
LMArena Expert1291980
GPQA Diamond47.5%—
Vectara Hallucination Rate5.1%—
MMLU68.7%—

Multimodal Not comparable

Mistral Small: 33.5 (#96), Qwen1.5 4b Chat: —

Multimodal benchmarks
BenchmarkMistral SmallQwen1.5 4b Chat
LMArena Vision1142—

Multilingual Mistral Small leads

Mistral Small: 45.5 (#169), Qwen1.5 4b Chat: 24.1 (#290)

Multilingual benchmarks
BenchmarkMistral SmallQwen1.5 4b Chat
LMArena Non-English1315979
LMArena Chinese13401024
LMArena German1340902
LMArena Russian1324952
LMArena French1337—
LMArena Japanese1275—
LMArena Korean1259—
LMArena Spanish1346—

Instruction Following Mistral Small leads

Mistral Small: 66.4 (#209), Qwen1.5 4b Chat: 49.0 (#300)

Instruction Following benchmarks
BenchmarkMistral SmallQwen1.5 4b Chat
LMArena Instruction Following1310978
LiveBench Instruction Following63.7%—

Long Context Mistral Small leads

Mistral Small: 40.4 (#156), Qwen1.5 4b Chat: 30.1 (#290)

Long Context benchmarks
BenchmarkMistral SmallQwen1.5 4b Chat
LMArena Longer Query1327988

Writing & Preference Mistral Small leads

Mistral Small: 52.5 (#171), Qwen1.5 4b Chat: 23.8 (#309)

Writing & Preference benchmarks
BenchmarkMistral SmallQwen1.5 4b Chat
LMArena Text1338997
LMArena Creative Writing1305969
LMArena Multi-Turn1344977
LiveBench Language30.5%—

Frequently asked questions

Is Mistral Small better than Qwen1.5 4b Chat?

Mistral Small is the stronger model overall, scoring 33.4 to 28.8 on the Noometry Index.

Is Mistral Small or Qwen1.5 4b Chat better for coding?

Mistral Small scores higher on coding benchmarks: 34.0 versus 29.1 in the Noometry coding category.

How many benchmarks do Mistral Small and Qwen1.5 4b Chat share?

13 benchmarks have published results for both models. Mistral Small has 39 scored results on Noometry and Qwen1.5 4b Chat has 13.

Related comparisons

Go deeper