Model comparison

Mercury vs Wizardlm 70b

Mercury is the stronger model overall, scoring 37.6 to 33.0 on the Noometry Index.

Last verified . 8 shared benchmarks.

Mercury Inception

37.6

Rank #199 Confirmed

Wizardlm 70b Microsoft

33.0

Rank #249 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Mercury scores higher in 5 categories and Wizardlm 70b in 1 category; 6 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Mercury leads 41.6 to 29.6.
  • Wizardlm 70b has downloadable open weights; the other is API-only.

Side by side

Mercury and Wizardlm 70b specifications
MercuryWizardlm 70b
ProviderInceptionMicrosoft
Noometry Index37.633.0
Released——
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked912

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury leads

Mercury: 38.7 (#170), Wizardlm 70b: 31.4 (#285)

Coding benchmarks
BenchmarkMercuryWizardlm 70b
LMArena Coding13221081

Reasoning Wizardlm 70b leads

Mercury: 17.5 (#293), Wizardlm 70b: 20.7 (#234)

Reasoning benchmarks
BenchmarkMercuryWizardlm 70b
LMArena Hard Prompts12851079
Kagi LLM Benchmark21.6%—

Math Not comparable

Mercury: —, Wizardlm 70b: 32.2 (#218)

Math benchmarks
BenchmarkMercuryWizardlm 70b
LMArena Math—1116

Multilingual Mercury leads

Mercury: 41.6 (#206), Wizardlm 70b: 29.6 (#265)

Multilingual benchmarks
BenchmarkMercuryWizardlm 70b
LMArena Non-English12601078
LMArena Chinese—1052
LMArena German—1083
LMArena Russian—1155

Instruction Following Mercury leads

Mercury: 65.2 (#224), Wizardlm 70b: 56.3 (#273)

Instruction Following benchmarks
BenchmarkMercuryWizardlm 70b
LMArena Instruction Following12391093

Long Context Mercury leads

Mercury: 38.4 (#198), Wizardlm 70b: 33.3 (#263)

Long Context benchmarks
BenchmarkMercuryWizardlm 70b
LMArena Longer Query12661097

Writing & Preference Mercury leads

Mercury: 46.2 (#221), Wizardlm 70b: 34.8 (#269)

Writing & Preference benchmarks
BenchmarkMercuryWizardlm 70b
LMArena Text12821120
LMArena Creative Writing11911149
LMArena Multi-Turn12821108

Frequently asked questions

Is Mercury better than Wizardlm 70b?

Mercury is the stronger model overall, scoring 37.6 to 33.0 on the Noometry Index.

Is Mercury or Wizardlm 70b better for coding?

Mercury scores higher on coding benchmarks: 38.7 versus 31.4 in the Noometry coding category.

How many benchmarks do Mercury and Wizardlm 70b share?

8 benchmarks have published results for both models. Mercury has 9 scored results on Noometry and Wizardlm 70b has 12.

Related comparisons

Go deeper