Model comparison

Mistral Small vs Qwen2.5 14B Instruct

Mistral Small has enough public results to be ranked (#243); Qwen2.5 14B Instruct does not yet, so treat this comparison as directional.

Last verified . 3 shared benchmarks.

Mistral Small Mistral AI

33.4

Rank #243 Confirmed

Summary

  • They share 3 benchmarks with published results for both. Mistral Small scores higher in 0 categories and Qwen2.5 14B Instruct in 1 category; one gap is clear of the uncertainty.
  • The widest gap is in coding, where Qwen2.5 14B Instruct leads 37.7 to 34.0.
  • The biggest single-benchmark swing is BigCodeBench Complete: 46.6% for Mistral Small and 52.2% for Qwen2.5 14B Instruct.
  • Mistral Small is cheaper at $0.15 / $0.60 per million input/output tokens, against $0.35 / $1.40 for Qwen2.5 14B Instruct.
  • Mistral Small accepts more context: 262K tokens versus 131K.

Side by side

Mistral Small and Qwen2.5 14B Instruct specifications
Mistral SmallQwen2.5 14B Instruct
ProviderMistral AIAlibaba (Qwen)
Noometry Index33.438.7
Released2024-02-262024-09
WeightsOpenOpen
Context window262K131K
Max output256K8K
Input $ / M tokens$0.15$0.35
Output $ / M tokens$0.60$1.40
Results tracked393

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 14B Instruct leads

Mistral Small: 34.0 (#247), Qwen2.5 14B Instruct: 37.7 (#191)

Coding benchmarks
BenchmarkMistral SmallQwen2.5 14B Instruct
BigCodeBench Instruct36.1%39.8%
BigCodeBench Complete46.6%52.2%
SciCode26.5%—
LiveBench Coding36.2%—
LMArena Coding1362—
ALE-Bench497.62—

Agentic & Tool Use Not comparable

Mistral Small: 28.1 (#93), Qwen2.5 14B Instruct: —

Agentic & Tool Use benchmarks
BenchmarkMistral SmallQwen2.5 14B Instruct
Berkeley Function Calling Leaderboard37.1%—

Reasoning Not comparable

Mistral Small: 19.8 (#250), Qwen2.5 14B Instruct: —

Reasoning benchmarks
BenchmarkMistral SmallQwen2.5 14B Instruct
Kagi LLM Benchmark37.8%—
CritPt0%—
LiveBench Reasoning44.8%—
LMArena Hard Prompts1335—
DTBench70.9%—
LiveBench Data Analysis53.7%—
LMCA20.6%—
LiveBench44%—

Math Not comparable

Mistral Small: 16.4 (#293), Qwen2.5 14B Instruct: —

Math benchmarks
BenchmarkMistral SmallQwen2.5 14B Instruct
OTIS Mock AIME 2024-20255.8%—
LiveBench Math39.9%—
LMArena Math1341—
MATH Level 546.8%—

Knowledge Not comparable

Mistral Small: 31.0 (#222), Qwen2.5 14B Instruct: —

Knowledge benchmarks
BenchmarkMistral SmallQwen2.5 14B Instruct
MMLU68.7%79.9%
GPQA Diamond47.5%—
Vectara Hallucination Rate5.1%—
LMArena Expert1291—

Multimodal Not comparable

Mistral Small: 33.5 (#96), Qwen2.5 14B Instruct: —

Multimodal benchmarks
BenchmarkMistral SmallQwen2.5 14B Instruct
LMArena Vision1142—

Multilingual Not comparable

Mistral Small: 45.5 (#169), Qwen2.5 14B Instruct: —

Multilingual benchmarks
BenchmarkMistral SmallQwen2.5 14B Instruct
LMArena Non-English1315—
LMArena Chinese1340—
LMArena French1337—
LMArena German1340—
LMArena Japanese1275—
LMArena Korean1259—
LMArena Russian1324—
LMArena Spanish1346—

Instruction Following Not comparable

Mistral Small: 66.4 (#209), Qwen2.5 14B Instruct: —

Instruction Following benchmarks
BenchmarkMistral SmallQwen2.5 14B Instruct
LiveBench Instruction Following63.7%—
LMArena Instruction Following1310—

Long Context Not comparable

Mistral Small: 40.4 (#156), Qwen2.5 14B Instruct: —

Long Context benchmarks
BenchmarkMistral SmallQwen2.5 14B Instruct
LMArena Longer Query1327—

Writing & Preference Not comparable

Mistral Small: 52.5 (#171), Qwen2.5 14B Instruct: —

Writing & Preference benchmarks
BenchmarkMistral SmallQwen2.5 14B Instruct
LMArena Text1338—
LMArena Creative Writing1305—
LMArena Multi-Turn1344—
LiveBench Language30.5%—

Frequently asked questions

Is Mistral Small better than Qwen2.5 14B Instruct?

Mistral Small has enough public results to be ranked (#243); Qwen2.5 14B Instruct does not yet, so treat this comparison as directional.

Which is cheaper, Mistral Small or Qwen2.5 14B Instruct?

Mistral Small is cheaper. It lists at $0.15 per million input tokens and $0.60 per million output tokens; Qwen2.5 14B Instruct lists at $0.35 and $1.40.

Is Mistral Small or Qwen2.5 14B Instruct better for coding?

Qwen2.5 14B Instruct scores higher on coding benchmarks: 37.7 versus 34.0 in the Noometry coding category.

Which has the bigger context window?

Mistral Small does, with 262K tokens against 131K.

How many benchmarks do Mistral Small and Qwen2.5 14B Instruct share?

3 benchmarks have published results for both models. Mistral Small has 39 scored results on Noometry and Qwen2.5 14B Instruct has 3.

Related comparisons

Go deeper