Model comparison

Mixtral 8x22B vs Wizardlm 70b

Wizardlm 70b is the stronger model overall, scoring 33.0 to 27.1 on the Noometry Index.

Last verified . 12 shared benchmarks.

Mixtral 8x22B Mistral AI

27.1

Rank #333 Confirmed

Wizardlm 70b Microsoft

33.0

Rank #249 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Mixtral 8x22B scores higher in 4 categories and Wizardlm 70b in 3 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where Wizardlm 70b leads 32.2 to 22.9.

Side by side

Mixtral 8x22B and Wizardlm 70b specifications
Mixtral 8x22BWizardlm 70b
ProviderMistral AIMicrosoft
Noometry Index27.133.0
Released2024-04-17—
WeightsOpenOpen
Context window64K—
Max output64K—
Input $ / M tokens$2—
Output $ / M tokens$6—
Results tracked3412

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Wizardlm 70b leads

Mixtral 8x22B: 24.2 (#329), Wizardlm 70b: 31.4 (#285)

Coding benchmarks
BenchmarkMixtral 8x22BWizardlm 70b
LMArena Coding11661081
WeirdML3.2%—
BigCodeBench Instruct40.6%—
BigCodeBench Complete50.2%—
HumanEval+72%—
MBPP+64.3%—

Agentic & Tool Use Not comparable

Mixtral 8x22B: 23.1 (#127), Wizardlm 70b: —

Agentic & Tool Use benchmarks
BenchmarkMixtral 8x22BWizardlm 70b
Cybench7.5%—

Reasoning Too close to call

Mixtral 8x22B: 19.9 (#248), Wizardlm 70b: 20.7 (#234)

Reasoning benchmarks
BenchmarkMixtral 8x22BWizardlm 70b
LMArena Hard Prompts11501079
DTBench55.1%—
Epoch Capabilities Index122.03—
ForecastBench56.3—

Math Wizardlm 70b leads

Mixtral 8x22B: 22.9 (#275), Wizardlm 70b: 32.2 (#218)

Math benchmarks
BenchmarkMixtral 8x22BWizardlm 70b
LMArena Math11841116
Omni-MATH16.3%—
MATH Level 524.2%—

Knowledge Not comparable

Mixtral 8x22B: 15.1 (#293), Wizardlm 70b: —

Knowledge benchmarks
BenchmarkMixtral 8x22BWizardlm 70b
GPQA Diamond34.1%—
MMLU-Pro46%—
GPQA (HELM)33.4%—
LMArena Expert1113—
MMLU77.8%—

Multilingual Mixtral 8x22B leads

Mixtral 8x22B: 32.8 (#255), Wizardlm 70b: 29.6 (#265)

Multilingual benchmarks
BenchmarkMixtral 8x22BWizardlm 70b
LMArena Non-English11281078
LMArena Chinese11161052
LMArena German11411083
LMArena Russian11581155
LMArena French1166—
LMArena Japanese1037—
LMArena Korean1057—
LMArena Spanish1151—

Instruction Following Mixtral 8x22B leads

Mixtral 8x22B: 57.7 (#266), Wizardlm 70b: 56.3 (#273)

Instruction Following benchmarks
BenchmarkMixtral 8x22BWizardlm 70b
LMArena Instruction Following11471093
IFEval72.4%—

Long Context Mixtral 8x22B leads

Mixtral 8x22B: 34.7 (#247), Wizardlm 70b: 33.3 (#263)

Long Context benchmarks
BenchmarkMixtral 8x22BWizardlm 70b
LMArena Longer Query11441097

Writing & Preference Mixtral 8x22B leads

Mixtral 8x22B: 36.9 (#262), Wizardlm 70b: 34.8 (#269)

Writing & Preference benchmarks
BenchmarkMixtral 8x22BWizardlm 70b
LMArena Text11621120
LMArena Creative Writing11411149
LMArena Multi-Turn11301108
WildBench71.1%—

Frequently asked questions

Is Mixtral 8x22B better than Wizardlm 70b?

Wizardlm 70b is the stronger model overall, scoring 33.0 to 27.1 on the Noometry Index.

Is Mixtral 8x22B or Wizardlm 70b better for coding?

Wizardlm 70b scores higher on coding benchmarks: 31.4 versus 24.2 in the Noometry coding category.

How many benchmarks do Mixtral 8x22B and Wizardlm 70b share?

12 benchmarks have published results for both models. Mixtral 8x22B has 34 scored results on Noometry and Wizardlm 70b has 12.

Related comparisons

Go deeper