Model comparison

Mixtral 8x22B vs Wizardlm 13b

Wizardlm 13b is the stronger model overall, scoring 31.4 to 27.1 on the Noometry Index.

Last verified . 10 shared benchmarks.

Mixtral 8x22B Mistral AI

27.1

Rank #333 Confirmed

Wizardlm 13b Microsoft

31.4

Rank #274 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Mixtral 8x22B scores higher in 5 categories and Wizardlm 13b in 2 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where Wizardlm 13b leads 30.2 to 22.9.

Side by side

Mixtral 8x22B and Wizardlm 13b specifications
Mixtral 8x22BWizardlm 13b
ProviderMistral AIMicrosoft
Noometry Index27.131.4
Released2024-04-17—
WeightsOpenOpen
Context window64K—
Max output64K—
Input $ / M tokens$2—
Output $ / M tokens$6—
Results tracked3410

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Wizardlm 13b leads

Mixtral 8x22B: 24.2 (#329), Wizardlm 13b: 30.1 (#298)

Coding benchmarks
BenchmarkMixtral 8x22BWizardlm 13b
LMArena Coding11661035
WeirdML3.2%—
BigCodeBench Instruct40.6%—
BigCodeBench Complete50.2%—
HumanEval+72%—
MBPP+64.3%—

Agentic & Tool Use Not comparable

Mixtral 8x22B: 23.1 (#127), Wizardlm 13b: —

Agentic & Tool Use benchmarks
BenchmarkMixtral 8x22BWizardlm 13b
Cybench7.5%—

Reasoning Too close to call

Mixtral 8x22B: 19.9 (#248), Wizardlm 13b: 19.4 (#259)

Reasoning benchmarks
BenchmarkMixtral 8x22BWizardlm 13b
LMArena Hard Prompts11501018
DTBench55.1%—
Epoch Capabilities Index122.03—
ForecastBench56.3—

Math Wizardlm 13b leads

Mixtral 8x22B: 22.9 (#275), Wizardlm 13b: 30.2 (#238)

Math benchmarks
BenchmarkMixtral 8x22BWizardlm 13b
LMArena Math11841017
Omni-MATH16.3%—
MATH Level 524.2%—

Knowledge Not comparable

Mixtral 8x22B: 15.1 (#293), Wizardlm 13b: —

Knowledge benchmarks
BenchmarkMixtral 8x22BWizardlm 13b
GPQA Diamond34.1%—
MMLU-Pro46%—
GPQA (HELM)33.4%—
LMArena Expert1113—
MMLU77.8%—

Multilingual Mixtral 8x22B leads

Mixtral 8x22B: 32.8 (#255), Wizardlm 13b: 27.1 (#277)

Multilingual benchmarks
BenchmarkMixtral 8x22BWizardlm 13b
LMArena Non-English11281034
LMArena Chinese11161023
LMArena French1166—
LMArena German1141—
LMArena Japanese1037—
LMArena Korean1057—
LMArena Russian1158—
LMArena Spanish1151—

Instruction Following Mixtral 8x22B leads

Mixtral 8x22B: 57.7 (#266), Wizardlm 13b: 53.5 (#285)

Instruction Following benchmarks
BenchmarkMixtral 8x22BWizardlm 13b
LMArena Instruction Following11471048
IFEval72.4%—

Long Context Mixtral 8x22B leads

Mixtral 8x22B: 34.7 (#247), Wizardlm 13b: 32.0 (#273)

Long Context benchmarks
BenchmarkMixtral 8x22BWizardlm 13b
LMArena Longer Query11441054

Writing & Preference Mixtral 8x22B leads

Mixtral 8x22B: 36.9 (#262), Wizardlm 13b: 30.5 (#287)

Writing & Preference benchmarks
BenchmarkMixtral 8x22BWizardlm 13b
LMArena Text11621077
LMArena Creative Writing11411091
LMArena Multi-Turn11301047
WildBench71.1%—

Frequently asked questions

Is Mixtral 8x22B better than Wizardlm 13b?

Wizardlm 13b is the stronger model overall, scoring 31.4 to 27.1 on the Noometry Index.

Is Mixtral 8x22B or Wizardlm 13b better for coding?

Wizardlm 13b scores higher on coding benchmarks: 30.1 versus 24.2 in the Noometry coding category.

How many benchmarks do Mixtral 8x22B and Wizardlm 13b share?

10 benchmarks have published results for both models. Mixtral 8x22B has 34 scored results on Noometry and Wizardlm 13b has 10.

Related comparisons

Go deeper