Model comparison

Mistral vs Qwen2-72B

Mistral and Qwen2-72B score almost the same on the Noometry Index (29.9 vs 30.0), so choose on price, context window or the category you care about most.

Last verified . 17 shared benchmarks.

Mistral Mistral AI

29.9

Rank #303 Confirmed

Qwen2-72B Alibaba (Qwen)

30.0

Rank #300 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Mistral scores higher in 1 category and Qwen2-72B in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Qwen2-72B leads 61.7 to 52.6.
  • Qwen2-72B has downloadable open weights; the other is API-only.

Side by side

Mistral and Qwen2-72B specifications
MistralQwen2-72B
ProviderMistral AIAlibaba (Qwen)
Noometry Index29.930.0
Released—2024-06-07
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2226

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral leads

Mistral: 33.8 (#250), Qwen2-72B: 29.1 (#310)

Coding benchmarks
BenchmarkMistralQwen2-72B
LMArena Coding11621196
WeirdML—11.3%
BigCodeBench Instruct—38.5%
BigCodeBench Complete—54%

Agentic & Tool Use Not comparable

Mistral: —, Qwen2-72B: 17.0 (#146)

Agentic & Tool Use benchmarks
BenchmarkMistralQwen2-72B
TheAgentCompany—1.1%
METR Time Horizons—29.9%

Reasoning Too close to call

Mistral: 22.2 (#200), Qwen2-72B: 23.2 (#181)

Reasoning benchmarks
BenchmarkMistralQwen2-72B
LMArena Hard Prompts11491191
Epoch Capabilities Index—125.28

Math Qwen2-72B leads

Mistral: 22.3 (#278), Qwen2-72B: 30.2 (#236)

Math benchmarks
BenchmarkMistralQwen2-72B
LMArena Math11801235
Omni-MATH7.2%—
MATH Level 5—39.1%

Knowledge Qwen2-72B leads

Mistral: 16.6 (#288), Qwen2-72B: 21.2 (#275)

Knowledge benchmarks
BenchmarkMistralQwen2-72B
LMArena Expert11251171
GPQA Diamond—40.8%
MMLU-Pro27.7%—
GPQA (HELM)30.3%—
MMLU—82.4%

Multilingual Qwen2-72B leads

Mistral: 32.8 (#254), Qwen2-72B: 35.9 (#244)

Multilingual benchmarks
BenchmarkMistralQwen2-72B
LMArena Non-English11291176
LMArena Chinese11091240
LMArena French11801170
LMArena German11551151
LMArena Japanese10131111
LMArena Korean10321083
LMArena Russian11681169
LMArena Spanish11431169

Instruction Following Qwen2-72B leads

Mistral: 52.6 (#288), Qwen2-72B: 61.7 (#241)

Instruction Following benchmarks
BenchmarkMistralQwen2-72B
LMArena Instruction Following11521181
IFEval56.8%—

Long Context Qwen2-72B leads

Mistral: 35.0 (#245), Qwen2-72B: 36.1 (#235)

Long Context benchmarks
BenchmarkMistralQwen2-72B
LMArena Longer Query11531192

Writing & Preference Qwen2-72B leads

Mistral: 37.0 (#260), Qwen2-72B: 40.8 (#241)

Writing & Preference benchmarks
BenchmarkMistralQwen2-72B
LMArena Text11651203
LMArena Creative Writing11581181
LMArena Multi-Turn11471196
WildBench66%—

Frequently asked questions

Is Mistral better than Qwen2-72B?

Mistral and Qwen2-72B score almost the same on the Noometry Index (29.9 vs 30.0), so choose on price, context window or the category you care about most.

Is Mistral or Qwen2-72B better for coding?

Mistral scores higher on coding benchmarks: 33.8 versus 29.1 in the Noometry coding category.

How many benchmarks do Mistral and Qwen2-72B share?

17 benchmarks have published results for both models. Mistral has 22 scored results on Noometry and Qwen2-72B has 26.

Related comparisons

Go deeper