Model comparison

Mercury 2.5 vs Wizardlm 13b

Mercury 2.5 is the stronger model overall, scoring 33.5 to 31.4 on the Noometry Index.

Last verified . 0 shared benchmarks.

Mercury 2.5 Inception

33.5

Rank #242 Reported

Wizardlm 13b Microsoft

31.4

Rank #274 Confirmed

Summary

  • The widest gap is in coding, where Mercury 2.5 leads 39.5 to 30.1.
  • Wizardlm 13b has downloadable open weights; the other is API-only.

Side by side

Mercury 2.5 and Wizardlm 13b specifications
Mercury 2.5Wizardlm 13b
ProviderInceptionMicrosoft
Noometry Index33.531.4
Released2026-09-08—
WeightsProprietaryOpen
Context window260K—
Max output66K—
Input $ / M tokens$0.04—
Output $ / M tokens$0.15—
Results tracked410

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury 2.5 leads

Mercury 2.5: 39.5 (#156), Wizardlm 13b: 30.1 (#298)

Coding benchmarks
BenchmarkMercury 2.5Wizardlm 13b
SciCode38.5%—
LMArena Coding—1035
ALE-Bench301.65—

Reasoning Mercury 2.5 leads

Mercury 2.5: 22.4 (#193), Wizardlm 13b: 19.4 (#259)

Reasoning benchmarks
BenchmarkMercury 2.5Wizardlm 13b
CritPt0%—
LMArena Hard Prompts—1018

Math Wizardlm 13b leads

Mercury 2.5: 23.3 (#272), Wizardlm 13b: 30.2 (#238)

Math benchmarks
BenchmarkMercury 2.5Wizardlm 13b
ProofBench3%—
LMArena Math—1017

Multilingual Not comparable

Mercury 2.5: —, Wizardlm 13b: 27.1 (#277)

Multilingual benchmarks
BenchmarkMercury 2.5Wizardlm 13b
LMArena Non-English—1034
LMArena Chinese—1023

Instruction Following Not comparable

Mercury 2.5: —, Wizardlm 13b: 53.5 (#285)

Instruction Following benchmarks
BenchmarkMercury 2.5Wizardlm 13b
LMArena Instruction Following—1048

Long Context Not comparable

Mercury 2.5: —, Wizardlm 13b: 32.0 (#273)

Long Context benchmarks
BenchmarkMercury 2.5Wizardlm 13b
LMArena Longer Query—1054

Writing & Preference Not comparable

Mercury 2.5: —, Wizardlm 13b: 30.5 (#287)

Writing & Preference benchmarks
BenchmarkMercury 2.5Wizardlm 13b
LMArena Text—1077
LMArena Creative Writing—1091
LMArena Multi-Turn—1047

Frequently asked questions

Is Mercury 2.5 better than Wizardlm 13b?

Mercury 2.5 is the stronger model overall, scoring 33.5 to 31.4 on the Noometry Index.

Is Mercury 2.5 or Wizardlm 13b better for coding?

Mercury 2.5 scores higher on coding benchmarks: 39.5 versus 30.1 in the Noometry coding category.

How many benchmarks do Mercury 2.5 and Wizardlm 13b share?

0 benchmarks have published results for both models. Mercury 2.5 has 4 scored results on Noometry and Wizardlm 13b has 10.

Related comparisons

Go deeper