Model comparison

Mistral Large vs Wizardlm 13b

Mistral Large and Wizardlm 13b score almost the same on the Noometry Index (31.9 vs 31.4), so choose on price, context window or the category you care about most.

Last verified . 10 shared benchmarks.

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Wizardlm 13b Microsoft

31.4

Rank #274 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Mistral Large scores higher in 5 categories and Wizardlm 13b in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Mistral Large leads 67.9 to 53.5.

Side by side

Mistral Large and Wizardlm 13b specifications
Mistral LargeWizardlm 13b
ProviderMistral AIMicrosoft
Noometry Index31.931.4
Released2024-02-26—
WeightsOpenOpen
Context window131K—
Max output16K—
Input $ / M tokens$2—
Output $ / M tokens$6—
Results tracked5110

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Large leads

Mistral Large: 34.3 (#240), Wizardlm 13b: 30.1 (#298)

Coding benchmarks
BenchmarkMistral LargeWizardlm 13b
LMArena Coding12771035
SciCode36.2%—
BigCodeBench Instruct30%—
LiveBench Coding47.1%—
BigCodeBench Complete38.3%—
ALE-Bench264.7—
HumanEval+62.2%—
MBPP+59.5%—

Agentic & Tool Use Not comparable

Mistral Large: 28.6 (#89), Wizardlm 13b: —

Agentic & Tool Use benchmarks
BenchmarkMistral LargeWizardlm 13b
Berkeley Function Calling Leaderboard38.4%—

Reasoning Wizardlm 13b leads

Mistral Large: 15.8 (#310), Wizardlm 13b: 19.4 (#259)

Reasoning benchmarks
BenchmarkMistral LargeWizardlm 13b
LMArena Hard Prompts12571018
SimpleBench22.5%—
CritPt0%—
LiveBench Reasoning43.5%—
DTBench65.1%—
LiveBench Data Analysis50.1%—
LMCA16.7%—
Epoch Capabilities Index128.52—
ForecastBench57.1—
LiveBench48.4%—

Math Wizardlm 13b leads

Mistral Large: 18.2 (#291), Wizardlm 13b: 30.2 (#238)

Math benchmarks
BenchmarkMistral LargeWizardlm 13b
LMArena Math12621017
OTIS Mock AIME 2024-20258.5%—
Omni-MATH28.1%—
LiveBench Math42.5%—
MATH Level 550.3%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Not comparable

Mistral Large: 30.1 (#230), Wizardlm 13b: —

Knowledge benchmarks
BenchmarkMistral LargeWizardlm 13b
GPQA Diamond51.3%—
MMLU-Pro59.9%—
Confabulations21.4%—
Vectara Hallucination Rate4.5%—
GPQA (HELM)43.5%—
LMArena Expert1232—
MMLU80%—

Multilingual Mistral Large leads

Mistral Large: 40.0 (#219), Wizardlm 13b: 27.1 (#277)

Multilingual benchmarks
BenchmarkMistral LargeWizardlm 13b
LMArena Non-English12371034
LMArena Chinese12401023
LMArena French1325—
LMArena German1254—
LMArena Japanese1188—
LMArena Korean1202—
LMArena Russian1257—
LMArena Spanish1268—

Instruction Following Mistral Large leads

Mistral Large: 67.9 (#191), Wizardlm 13b: 53.5 (#285)

Instruction Following benchmarks
BenchmarkMistral LargeWizardlm 13b
LMArena Instruction Following12491048
LiveBench Instruction Following67.9%—
IFEval87.7%—

Long Context Mistral Large leads

Mistral Large: 38.3 (#199), Wizardlm 13b: 32.0 (#273)

Long Context benchmarks
BenchmarkMistral LargeWizardlm 13b
LMArena Longer Query12611054

Writing & Preference Mistral Large leads

Mistral Large: 40.7 (#242), Wizardlm 13b: 30.5 (#287)

Writing & Preference benchmarks
BenchmarkMistral LargeWizardlm 13b
LMArena Text12661077
LMArena Creative Writing12431091
LMArena Multi-Turn12601047
Short-Story Creative Writing69%—
EQ-Bench Creative Writing985—
WildBench80.1%—
LiveBench Language39.4%—

Frequently asked questions

Is Mistral Large better than Wizardlm 13b?

Mistral Large and Wizardlm 13b score almost the same on the Noometry Index (31.9 vs 31.4), so choose on price, context window or the category you care about most.

Is Mistral Large or Wizardlm 13b better for coding?

Mistral Large scores higher on coding benchmarks: 34.3 versus 30.1 in the Noometry coding category.

How many benchmarks do Mistral Large and Wizardlm 13b share?

10 benchmarks have published results for both models. Mistral Large has 51 scored results on Noometry and Wizardlm 13b has 10.

Related comparisons

Go deeper