Model comparison

Mistral 7B vs Wizardlm 13b

Wizardlm 13b is the stronger model overall, scoring 31.4 to 23.0 on the Noometry Index.

Last verified . 10 shared benchmarks.

Mistral 7B Mistral AI

23.0

Rank #351 Confirmed

Wizardlm 13b Microsoft

31.4

Rank #274 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Mistral 7B scores higher in 3 categories and Wizardlm 13b in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in math, where Wizardlm 13b leads 30.2 to 8.1.

Side by side

Mistral 7B and Wizardlm 13b specifications
Mistral 7BWizardlm 13b
ProviderMistral AIMicrosoft
Noometry Index23.031.4
Released2023-09-27—
WeightsOpenOpen
Context window8K—
Max output8K—
Input $ / M tokens$0.25—
Output $ / M tokens$0.25—
Results tracked3710

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Wizardlm 13b leads

Mistral 7B: 26.4 (#326), Wizardlm 13b: 30.1 (#298)

Coding benchmarks
BenchmarkMistral 7BWizardlm 13b
LMArena Coding10821035
BigCodeBench Instruct19.5%—
BigCodeBench Complete27.3%—
HumanEval+36%—
MBPP+42.1%—

Reasoning Wizardlm 13b leads

Mistral 7B: 13.1 (#336), Wizardlm 13b: 19.4 (#259)

Reasoning benchmarks
BenchmarkMistral 7BWizardlm 13b
LMArena Hard Prompts10671018
Chess Puzzles0%—
DTBench42.5%—
Adversarial NLI47.1%—
BIG-Bench Hard56.1%—
Epoch Capabilities Index112.21—
HellaSwag81%—
PIQA83%—
WinoGrande75.3%—

Math Wizardlm 13b leads

Mistral 7B: 8.1 (#325), Wizardlm 13b: 30.2 (#238)

Math benchmarks
BenchmarkMistral 7BWizardlm 13b
LMArena Math10851017
OTIS Mock AIME 2024-20250.3%—
MATH Level 53.7%—
GSM8K54.4%—

Knowledge Not comparable

Mistral 7B: 7.4 (#311), Wizardlm 13b: —

Knowledge benchmarks
BenchmarkMistral 7BWizardlm 13b
GPQA Diamond15.2%—
LMArena Expert1036—
ARC (AI2) Challenge78.6%—
BoolQ87.4%—
MMLU62.5%—
OpenBookQA79.8%—
TriviaQA75.2%—

Multilingual Wizardlm 13b leads

Mistral 7B: 25.8 (#283), Wizardlm 13b: 27.1 (#277)

Multilingual benchmarks
BenchmarkMistral 7BWizardlm 13b
LMArena Non-English10121034
LMArena Chinese10091023
LMArena French1037—
LMArena German987—
LMArena Japanese878—
LMArena Russian1018—
LMArena Spanish1026—

Instruction Following Too close to call

Mistral 7B: 54.2 (#280), Wizardlm 13b: 53.5 (#285)

Instruction Following benchmarks
BenchmarkMistral 7BWizardlm 13b
LMArena Instruction Following10601048

Long Context Too close to call

Mistral 7B: 32.2 (#271), Wizardlm 13b: 32.0 (#273)

Long Context benchmarks
BenchmarkMistral 7BWizardlm 13b
LMArena Longer Query10601054

Writing & Preference Too close to call

Mistral 7B: 30.7 (#286), Wizardlm 13b: 30.5 (#287)

Writing & Preference benchmarks
BenchmarkMistral 7BWizardlm 13b
LMArena Text10901077
LMArena Creative Writing10681091
LMArena Multi-Turn10621047

Frequently asked questions

Is Mistral 7B better than Wizardlm 13b?

Wizardlm 13b is the stronger model overall, scoring 31.4 to 23.0 on the Noometry Index.

Is Mistral 7B or Wizardlm 13b better for coding?

Wizardlm 13b scores higher on coding benchmarks: 30.1 versus 26.4 in the Noometry coding category.

How many benchmarks do Mistral 7B and Wizardlm 13b share?

10 benchmarks have published results for both models. Mistral 7B has 37 scored results on Noometry and Wizardlm 13b has 10.

Related comparisons

Go deeper