Model comparison

Gemma 7B vs Wizardlm 13b

Wizardlm 13b is the stronger model overall, scoring 31.4 to 30.0 on the Noometry Index.

Last verified . 10 shared benchmarks.

Gemma 7B Google

30.0

Rank #299 Confirmed

Wizardlm 13b Microsoft

31.4

Rank #274 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Gemma 7B scores higher in 3 categories and Wizardlm 13b in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Wizardlm 13b leads 30.5 to 27.1.

Side by side

Gemma 7B and Wizardlm 13b specifications
Gemma 7BWizardlm 13b
ProviderGoogleMicrosoft
Noometry Index30.031.4
Released2024-02-21—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked2710

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Gemma 7B: 30.5 (#294), Wizardlm 13b: 30.1 (#298)

Coding benchmarks
BenchmarkGemma 7BWizardlm 13b
LMArena Coding10481035
HumanEval+28.7%—
MBPP+43.4%—

Reasoning Too close to call

Gemma 7B: 19.9 (#249), Wizardlm 13b: 19.4 (#259)

Reasoning benchmarks
BenchmarkGemma 7BWizardlm 13b
LMArena Hard Prompts10421018
Adversarial NLI48.7%—
BIG-Bench Hard55.1%—
Epoch Capabilities Index111.99—
HellaSwag82.2%—
PIQA81.2%—
WinoGrande79%—

Math Gemma 7B leads

Gemma 7B: 31.2 (#228), Wizardlm 13b: 30.2 (#238)

Math benchmarks
BenchmarkGemma 7BWizardlm 13b
LMArena Math10661017
GSM8K46.4%—

Knowledge Not comparable

Gemma 7B: 27.3 (#252), Wizardlm 13b: —

Knowledge benchmarks
BenchmarkGemma 7BWizardlm 13b
LMArena Expert1001—
ARC (AI2) Challenge78.3%—
BoolQ83.2%—
MMLU66.1%—
OpenBookQA78.6%—
TriviaQA72.3%—

Multilingual Wizardlm 13b leads

Gemma 7B: 25.1 (#287), Wizardlm 13b: 27.1 (#277)

Multilingual benchmarks
BenchmarkGemma 7BWizardlm 13b
LMArena Non-English9991034
LMArena Chinese10351023
LMArena French1025—
LMArena Russian993—

Instruction Following Wizardlm 13b leads

Gemma 7B: 51.5 (#295), Wizardlm 13b: 53.5 (#285)

Instruction Following benchmarks
BenchmarkGemma 7BWizardlm 13b
LMArena Instruction Following10171048

Long Context Too close to call

Gemma 7B: 31.1 (#282), Wizardlm 13b: 32.0 (#273)

Long Context benchmarks
BenchmarkGemma 7BWizardlm 13b
LMArena Longer Query10221054

Writing & Preference Wizardlm 13b leads

Gemma 7B: 27.1 (#302), Wizardlm 13b: 30.5 (#287)

Writing & Preference benchmarks
BenchmarkGemma 7BWizardlm 13b
LMArena Text10561077
LMArena Creative Writing10241091
LMArena Multi-Turn9631047

Frequently asked questions

Is Gemma 7B better than Wizardlm 13b?

Wizardlm 13b is the stronger model overall, scoring 31.4 to 30.0 on the Noometry Index.

Is Gemma 7B or Wizardlm 13b better for coding?

They score almost the same on coding (30.5 vs 30.1); test both on your own repository before choosing.

How many benchmarks do Gemma 7B and Wizardlm 13b share?

10 benchmarks have published results for both models. Gemma 7B has 27 scored results on Noometry and Wizardlm 13b has 10.

Related comparisons

Go deeper