Model comparison

Llama 3.1-405B vs Wizardlm 13b

Llama 3.1-405B and Wizardlm 13b score almost the same on the Noometry Index (30.7 vs 31.4), so choose on price, context window or the category you care about most.

Last verified . 10 shared benchmarks.

Llama 3.1-405B Meta

30.7

Rank #288 Confirmed

Wizardlm 13b Microsoft

31.4

Rank #274 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Llama 3.1-405B scores higher in 5 categories and Wizardlm 13b in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Llama 3.1-405B leads 40.7 to 27.1.

Side by side

Llama 3.1-405B and Wizardlm 13b specifications
Llama 3.1-405BWizardlm 13b
ProviderMetaMicrosoft
Noometry Index30.731.4
Released2024-07-23—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4210

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Llama 3.1-405B leads

Llama 3.1-405B: 33.1 (#262), Wizardlm 13b: 30.1 (#298)

Coding benchmarks
BenchmarkLlama 3.1-405BWizardlm 13b
LMArena Coding12911035
WeirdML21.4%—

Agentic & Tool Use Not comparable

Llama 3.1-405B: 21.0 (#140), Wizardlm 13b: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.1-405BWizardlm 13b
TheAgentCompany7.4%—
Cybench7.5%—

Reasoning Wizardlm 13b leads

Llama 3.1-405B: 16.8 (#300), Wizardlm 13b: 19.4 (#259)

Reasoning benchmarks
BenchmarkLlama 3.1-405BWizardlm 13b
LMArena Hard Prompts12691018
SimpleBench23%—
Kagi LLM Benchmark45%—
DTBench61.4%—
BIG-Bench Hard82.9%—
Epoch Capabilities Index128.75—
ForecastBench59.9—
HellaSwag89.2%—
PIQA85.9%—
WinoGrande89.2%—

Math Wizardlm 13b leads

Llama 3.1-405B: 18.4 (#290), Wizardlm 13b: 30.2 (#238)

Math benchmarks
BenchmarkLlama 3.1-405BWizardlm 13b
LMArena Math12811017
OTIS Mock AIME 2024-20259.7%—
Omni-MATH24.9%—
MATH Level 549.8%—

Knowledge Not comparable

Llama 3.1-405B: 30.4 (#227), Wizardlm 13b: —

Knowledge benchmarks
BenchmarkLlama 3.1-405BWizardlm 13b
GPQA Diamond50.9%—
MMLU-Pro72.3%—
Confabulations17.6%—
GPQA (HELM)52.2%—
LMArena Expert1243—
ARC (AI2) Challenge95.3%—
MMLU84.5%—
TriviaQA82.7%—

Multilingual Llama 3.1-405B leads

Llama 3.1-405B: 40.7 (#214), Wizardlm 13b: 27.1 (#277)

Multilingual benchmarks
BenchmarkLlama 3.1-405BWizardlm 13b
LMArena Non-English12481034
LMArena Chinese12421023
LMArena French1279—
LMArena German1252—
LMArena Japanese1208—
LMArena Korean1184—
LMArena Russian1265—
LMArena Spanish1260—

Instruction Following Llama 3.1-405B leads

Llama 3.1-405B: 65.9 (#214), Wizardlm 13b: 53.5 (#285)

Instruction Following benchmarks
BenchmarkLlama 3.1-405BWizardlm 13b
LMArena Instruction Following12591048
IFEval81.1%—

Long Context Llama 3.1-405B leads

Llama 3.1-405B: 38.4 (#197), Wizardlm 13b: 32.0 (#273)

Long Context benchmarks
BenchmarkLlama 3.1-405BWizardlm 13b
LMArena Longer Query12661054

Writing & Preference Llama 3.1-405B leads

Llama 3.1-405B: 38.9 (#251), Wizardlm 13b: 30.5 (#287)

Writing & Preference benchmarks
BenchmarkLlama 3.1-405BWizardlm 13b
LMArena Text12841077
LMArena Creative Writing12621091
LMArena Multi-Turn12971047
EQ-Bench Creative Writing870—
WildBench78.3%—

Frequently asked questions

Is Llama 3.1-405B better than Wizardlm 13b?

Llama 3.1-405B and Wizardlm 13b score almost the same on the Noometry Index (30.7 vs 31.4), so choose on price, context window or the category you care about most.

Is Llama 3.1-405B or Wizardlm 13b better for coding?

Llama 3.1-405B scores higher on coding benchmarks: 33.1 versus 30.1 in the Noometry coding category.

How many benchmarks do Llama 3.1-405B and Wizardlm 13b share?

10 benchmarks have published results for both models. Llama 3.1-405B has 42 scored results on Noometry and Wizardlm 13b has 10.

Related comparisons

Go deeper