Model comparison

Mistral vs Wizardlm 70b

Wizardlm 70b is the stronger model overall, scoring 33.0 to 29.9 on the Noometry Index.

Last verified . 12 shared benchmarks.

Mistral Mistral AI

29.9

Rank #303 Confirmed

Wizardlm 70b Microsoft

33.0

Rank #249 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Mistral scores higher in 5 categories and Wizardlm 70b in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Wizardlm 70b leads 32.2 to 22.3.
  • Wizardlm 70b has downloadable open weights; the other is API-only.

Side by side

Mistral and Wizardlm 70b specifications
MistralWizardlm 70b
ProviderMistral AIMicrosoft
Noometry Index29.933.0
Released——
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2212

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral leads

Mistral: 33.8 (#250), Wizardlm 70b: 31.4 (#285)

Coding benchmarks
BenchmarkMistralWizardlm 70b
LMArena Coding11621081

Reasoning Mistral leads

Mistral: 22.2 (#200), Wizardlm 70b: 20.7 (#234)

Reasoning benchmarks
BenchmarkMistralWizardlm 70b
LMArena Hard Prompts11491079

Math Wizardlm 70b leads

Mistral: 22.3 (#278), Wizardlm 70b: 32.2 (#218)

Math benchmarks
BenchmarkMistralWizardlm 70b
LMArena Math11801116
Omni-MATH7.2%—

Knowledge Not comparable

Mistral: 16.6 (#288), Wizardlm 70b: —

Knowledge benchmarks
BenchmarkMistralWizardlm 70b
MMLU-Pro27.7%—
GPQA (HELM)30.3%—
LMArena Expert1125—

Multilingual Mistral leads

Mistral: 32.8 (#254), Wizardlm 70b: 29.6 (#265)

Multilingual benchmarks
BenchmarkMistralWizardlm 70b
LMArena Non-English11291078
LMArena Chinese11091052
LMArena German11551083
LMArena Russian11681155
LMArena French1180—
LMArena Japanese1013—
LMArena Korean1032—
LMArena Spanish1143—

Instruction Following Wizardlm 70b leads

Mistral: 52.6 (#288), Wizardlm 70b: 56.3 (#273)

Instruction Following benchmarks
BenchmarkMistralWizardlm 70b
LMArena Instruction Following11521093
IFEval56.8%—

Long Context Mistral leads

Mistral: 35.0 (#245), Wizardlm 70b: 33.3 (#263)

Long Context benchmarks
BenchmarkMistralWizardlm 70b
LMArena Longer Query11531097

Writing & Preference Mistral leads

Mistral: 37.0 (#260), Wizardlm 70b: 34.8 (#269)

Writing & Preference benchmarks
BenchmarkMistralWizardlm 70b
LMArena Text11651120
LMArena Creative Writing11581149
LMArena Multi-Turn11471108
WildBench66%—

Frequently asked questions

Is Mistral better than Wizardlm 70b?

Wizardlm 70b is the stronger model overall, scoring 33.0 to 29.9 on the Noometry Index.

Is Mistral or Wizardlm 70b better for coding?

Mistral scores higher on coding benchmarks: 33.8 versus 31.4 in the Noometry coding category.

How many benchmarks do Mistral and Wizardlm 70b share?

12 benchmarks have published results for both models. Mistral has 22 scored results on Noometry and Wizardlm 70b has 12.

Related comparisons

Go deeper