Model comparison

Mistral Small vs Wizardlm 70b

Mistral Small and Wizardlm 70b score almost the same on the Noometry Index (33.4 vs 33.0), so choose on price, context window or the category you care about most.

Last verified . 12 shared benchmarks.

Mistral Small Mistral AI

33.4

Rank #243 Confirmed

Wizardlm 70b Microsoft

33.0

Rank #249 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Mistral Small scores higher in 5 categories and Wizardlm 70b in 2 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Small leads 52.5 to 34.8.

Side by side

Mistral Small and Wizardlm 70b specifications
Mistral SmallWizardlm 70b
ProviderMistral AIMicrosoft
Noometry Index33.433.0
Released2024-02-26—
WeightsOpenOpen
Context window262K—
Max output256K—
Input $ / M tokens$0.15—
Output $ / M tokens$0.60—
Results tracked3912

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Small leads

Mistral Small: 34.0 (#247), Wizardlm 70b: 31.4 (#285)

Coding benchmarks
BenchmarkMistral SmallWizardlm 70b
LMArena Coding13621081
SciCode26.5%—
BigCodeBench Instruct36.1%—
LiveBench Coding36.2%—
BigCodeBench Complete46.6%—
ALE-Bench497.62—

Agentic & Tool Use Not comparable

Mistral Small: 28.1 (#93), Wizardlm 70b: —

Agentic & Tool Use benchmarks
BenchmarkMistral SmallWizardlm 70b
Berkeley Function Calling Leaderboard37.1%—

Reasoning Too close to call

Mistral Small: 19.8 (#250), Wizardlm 70b: 20.7 (#234)

Reasoning benchmarks
BenchmarkMistral SmallWizardlm 70b
LMArena Hard Prompts13351079
Kagi LLM Benchmark37.8%—
CritPt0%—
LiveBench Reasoning44.8%—
DTBench70.9%—
LiveBench Data Analysis53.7%—
LMCA20.6%—
LiveBench44%—

Math Wizardlm 70b leads

Mistral Small: 16.4 (#293), Wizardlm 70b: 32.2 (#218)

Math benchmarks
BenchmarkMistral SmallWizardlm 70b
LMArena Math13411116
OTIS Mock AIME 2024-20255.8%—
LiveBench Math39.9%—
MATH Level 546.8%—

Knowledge Not comparable

Mistral Small: 31.0 (#222), Wizardlm 70b: —

Knowledge benchmarks
BenchmarkMistral SmallWizardlm 70b
GPQA Diamond47.5%—
Vectara Hallucination Rate5.1%—
LMArena Expert1291—
MMLU68.7%—

Multimodal Not comparable

Mistral Small: 33.5 (#96), Wizardlm 70b: —

Multimodal benchmarks
BenchmarkMistral SmallWizardlm 70b
LMArena Vision1142—

Multilingual Mistral Small leads

Mistral Small: 45.5 (#169), Wizardlm 70b: 29.6 (#265)

Multilingual benchmarks
BenchmarkMistral SmallWizardlm 70b
LMArena Non-English13151078
LMArena Chinese13401052
LMArena German13401083
LMArena Russian13241155
LMArena French1337—
LMArena Japanese1275—
LMArena Korean1259—
LMArena Spanish1346—

Instruction Following Mistral Small leads

Mistral Small: 66.4 (#209), Wizardlm 70b: 56.3 (#273)

Instruction Following benchmarks
BenchmarkMistral SmallWizardlm 70b
LMArena Instruction Following13101093
LiveBench Instruction Following63.7%—

Long Context Mistral Small leads

Mistral Small: 40.4 (#156), Wizardlm 70b: 33.3 (#263)

Long Context benchmarks
BenchmarkMistral SmallWizardlm 70b
LMArena Longer Query13271097

Writing & Preference Mistral Small leads

Mistral Small: 52.5 (#171), Wizardlm 70b: 34.8 (#269)

Writing & Preference benchmarks
BenchmarkMistral SmallWizardlm 70b
LMArena Text13381120
LMArena Creative Writing13051149
LMArena Multi-Turn13441108
LiveBench Language30.5%—

Frequently asked questions

Is Mistral Small better than Wizardlm 70b?

Mistral Small and Wizardlm 70b score almost the same on the Noometry Index (33.4 vs 33.0), so choose on price, context window or the category you care about most.

Is Mistral Small or Wizardlm 70b better for coding?

Mistral Small scores higher on coding benchmarks: 34.0 versus 31.4 in the Noometry coding category.

How many benchmarks do Mistral Small and Wizardlm 70b share?

12 benchmarks have published results for both models. Mistral Small has 39 scored results on Noometry and Wizardlm 70b has 12.

Related comparisons

Go deeper