Model comparison

Mercury 2.5 vs Wizardlm 70b

Mercury 2.5 and Wizardlm 70b score almost the same on the Noometry Index (33.5 vs 33.0), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

Mercury 2.5 Inception

33.5

Rank #242 Reported

Wizardlm 70b Microsoft

33.0

Rank #249 Confirmed

Summary

  • The widest gap is in math, where Wizardlm 70b leads 32.2 to 23.3.
  • Wizardlm 70b has downloadable open weights; the other is API-only.

Side by side

Mercury 2.5 and Wizardlm 70b specifications
Mercury 2.5Wizardlm 70b
ProviderInceptionMicrosoft
Noometry Index33.533.0
Released2026-09-08—
WeightsProprietaryOpen
Context window260K—
Max output66K—
Input $ / M tokens$0.04—
Output $ / M tokens$0.15—
Results tracked412

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury 2.5 leads

Mercury 2.5: 39.5 (#156), Wizardlm 70b: 31.4 (#285)

Coding benchmarks
BenchmarkMercury 2.5Wizardlm 70b
SciCode38.5%—
LMArena Coding—1081
ALE-Bench301.65—

Reasoning Mercury 2.5 leads

Mercury 2.5: 22.4 (#193), Wizardlm 70b: 20.7 (#234)

Reasoning benchmarks
BenchmarkMercury 2.5Wizardlm 70b
CritPt0%—
LMArena Hard Prompts—1079

Math Wizardlm 70b leads

Mercury 2.5: 23.3 (#272), Wizardlm 70b: 32.2 (#218)

Math benchmarks
BenchmarkMercury 2.5Wizardlm 70b
ProofBench3%—
LMArena Math—1116

Multilingual Not comparable

Mercury 2.5: —, Wizardlm 70b: 29.6 (#265)

Multilingual benchmarks
BenchmarkMercury 2.5Wizardlm 70b
LMArena Non-English—1078
LMArena Chinese—1052
LMArena German—1083
LMArena Russian—1155

Instruction Following Not comparable

Mercury 2.5: —, Wizardlm 70b: 56.3 (#273)

Instruction Following benchmarks
BenchmarkMercury 2.5Wizardlm 70b
LMArena Instruction Following—1093

Long Context Not comparable

Mercury 2.5: —, Wizardlm 70b: 33.3 (#263)

Long Context benchmarks
BenchmarkMercury 2.5Wizardlm 70b
LMArena Longer Query—1097

Writing & Preference Not comparable

Mercury 2.5: —, Wizardlm 70b: 34.8 (#269)

Writing & Preference benchmarks
BenchmarkMercury 2.5Wizardlm 70b
LMArena Text—1120
LMArena Creative Writing—1149
LMArena Multi-Turn—1108

Frequently asked questions

Is Mercury 2.5 better than Wizardlm 70b?

Mercury 2.5 and Wizardlm 70b score almost the same on the Noometry Index (33.5 vs 33.0), so choose on price, context window or the category you care about most.

Is Mercury 2.5 or Wizardlm 70b better for coding?

Mercury 2.5 scores higher on coding benchmarks: 39.5 versus 31.4 in the Noometry coding category.

How many benchmarks do Mercury 2.5 and Wizardlm 70b share?

0 benchmarks have published results for both models. Mercury 2.5 has 4 scored results on Noometry and Wizardlm 70b has 12.

Related comparisons

Go deeper