Model comparison

Mistral Small 3.1 vs Wizardlm 13b

Mistral Small 3.1 and Wizardlm 13b score almost the same on the Noometry Index (31.7 vs 31.4), so choose on price, context window or the category you care about most.

Last verified . 10 shared benchmarks.

Mistral Small 3.1 Mistral AI

31.7

Rank #269 Confirmed

Wizardlm 13b Microsoft

31.4

Rank #274 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Mistral Small 3.1 scores higher in 6 categories and Wizardlm 13b in 1 category; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where Wizardlm 13b leads 30.2 to 14.7.

Side by side

Mistral Small 3.1 and Wizardlm 13b specifications
Mistral Small 3.1Wizardlm 13b
ProviderMistral AIMicrosoft
Noometry Index31.731.4
Released2025-03-17—
WeightsOpenOpen
Context window128K—
Max output102K—
Input $ / M tokens$0.35—
Output $ / M tokens$0.56—
Results tracked2810

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Small 3.1 leads

Mistral Small 3.1: 38.3 (#179), Wizardlm 13b: 30.1 (#298)

Coding benchmarks
BenchmarkMistral Small 3.1Wizardlm 13b
LMArena Coding13091035

Reasoning Too close to call

Mistral Small 3.1: 19.7 (#254), Wizardlm 13b: 19.4 (#259)

Reasoning benchmarks
BenchmarkMistral Small 3.1Wizardlm 13b
LMArena Hard Prompts12781018
Chess Puzzles1%—
Epoch Capabilities Index127.48—

Math Wizardlm 13b leads

Mistral Small 3.1: 14.7 (#301), Wizardlm 13b: 30.2 (#238)

Math benchmarks
BenchmarkMistral Small 3.1Wizardlm 13b
LMArena Math12621017
OTIS Mock AIME 2024-20253.9%—
Omni-MATH24.8%—

Knowledge Not comparable

Mistral Small 3.1: 22.6 (#271), Wizardlm 13b: —

Knowledge benchmarks
BenchmarkMistral Small 3.1Wizardlm 13b
GPQA Diamond41.9%—
MMLU-Pro61%—
GPQA (HELM)39.2%—
LMArena Expert1257—

Multimodal Not comparable

Mistral Small 3.1: 33.2 (#99), Wizardlm 13b: —

Multimodal benchmarks
BenchmarkMistral Small 3.1Wizardlm 13b
LMArena Vision1136—

Multilingual Mistral Small 3.1 leads

Mistral Small 3.1: 41.2 (#209), Wizardlm 13b: 27.1 (#277)

Multilingual benchmarks
BenchmarkMistral Small 3.1Wizardlm 13b
LMArena Non-English12551034
LMArena Chinese12531023
LMArena French1273—
LMArena German1266—
LMArena Japanese1208—
LMArena Korean1206—
LMArena Russian1263—
LMArena Spanish1283—

Instruction Following Mistral Small 3.1 leads

Mistral Small 3.1: 63.6 (#230), Wizardlm 13b: 53.5 (#285)

Instruction Following benchmarks
BenchmarkMistral Small 3.1Wizardlm 13b
LMArena Instruction Following12641048
IFEval75%—

Long Context Mistral Small 3.1 leads

Mistral Small 3.1: 39.5 (#178), Wizardlm 13b: 32.0 (#273)

Long Context benchmarks
BenchmarkMistral Small 3.1Wizardlm 13b
LMArena Longer Query12991054

Writing & Preference Mistral Small 3.1 leads

Mistral Small 3.1: 37.0 (#259), Wizardlm 13b: 30.5 (#287)

Writing & Preference benchmarks
BenchmarkMistral Small 3.1Wizardlm 13b
LMArena Text12771077
LMArena Creative Writing12531091
LMArena Multi-Turn12701047
EQ-Bench Creative Writing761—
WildBench78.8%—

Frequently asked questions

Is Mistral Small 3.1 better than Wizardlm 13b?

Mistral Small 3.1 and Wizardlm 13b score almost the same on the Noometry Index (31.7 vs 31.4), so choose on price, context window or the category you care about most.

Is Mistral Small 3.1 or Wizardlm 13b better for coding?

Mistral Small 3.1 scores higher on coding benchmarks: 38.3 versus 30.1 in the Noometry coding category.

How many benchmarks do Mistral Small 3.1 and Wizardlm 13b share?

10 benchmarks have published results for both models. Mistral Small 3.1 has 28 scored results on Noometry and Wizardlm 13b has 10.

Related comparisons

Go deeper