Model comparison

Qwen2.5 Plus 1127 vs Wizardlm 13b

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 31.4 on the Noometry Index.

Last verified . 10 shared benchmarks.

Qwen2.5 Plus 1127 Alibaba (Qwen)

38.8

Rank #181 Confirmed

Wizardlm 13b Microsoft

31.4

Rank #274 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Qwen2.5 Plus 1127 scores higher in 7 categories and Wizardlm 13b in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen2.5 Plus 1127 leads 49.4 to 30.5.
  • Wizardlm 13b has downloadable open weights; the other is API-only.

Side by side

Qwen2.5 Plus 1127 and Wizardlm 13b specifications
Qwen2.5 Plus 1127Wizardlm 13b
ProviderAlibaba (Qwen)Microsoft
Noometry Index38.831.4
Released——
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1410

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 Plus 1127 leads

Qwen2.5 Plus 1127: 38.5 (#175), Wizardlm 13b: 30.1 (#298)

Coding benchmarks
BenchmarkQwen2.5 Plus 1127Wizardlm 13b
LMArena Coding13141035

Reasoning Qwen2.5 Plus 1127 leads

Qwen2.5 Plus 1127: 25.9 (#141), Wizardlm 13b: 19.4 (#259)

Reasoning benchmarks
BenchmarkQwen2.5 Plus 1127Wizardlm 13b
LMArena Hard Prompts12991018

Math Qwen2.5 Plus 1127 leads

Qwen2.5 Plus 1127: 36.1 (#174), Wizardlm 13b: 30.2 (#238)

Math benchmarks
BenchmarkQwen2.5 Plus 1127Wizardlm 13b
LMArena Math12981017

Knowledge Not comparable

Qwen2.5 Plus 1127: 35.5 (#183), Wizardlm 13b: —

Knowledge benchmarks
BenchmarkQwen2.5 Plus 1127Wizardlm 13b
LMArena Expert1289—

Multilingual Qwen2.5 Plus 1127 leads

Qwen2.5 Plus 1127: 41.9 (#201), Wizardlm 13b: 27.1 (#277)

Multilingual benchmarks
BenchmarkQwen2.5 Plus 1127Wizardlm 13b
LMArena Non-English12651034
LMArena Chinese13141023
LMArena German1231—
LMArena Japanese1207—
LMArena Russian1271—

Instruction Following Qwen2.5 Plus 1127 leads

Qwen2.5 Plus 1127: 67.2 (#199), Wizardlm 13b: 53.5 (#285)

Instruction Following benchmarks
BenchmarkQwen2.5 Plus 1127Wizardlm 13b
LMArena Instruction Following12751048

Long Context Qwen2.5 Plus 1127 leads

Qwen2.5 Plus 1127: 39.2 (#184), Wizardlm 13b: 32.0 (#273)

Long Context benchmarks
BenchmarkQwen2.5 Plus 1127Wizardlm 13b
LMArena Longer Query12921054

Writing & Preference Qwen2.5 Plus 1127 leads

Qwen2.5 Plus 1127: 49.4 (#192), Wizardlm 13b: 30.5 (#287)

Writing & Preference benchmarks
BenchmarkQwen2.5 Plus 1127Wizardlm 13b
LMArena Text12991077
LMArena Creative Writing12621091
LMArena Multi-Turn12991047

Frequently asked questions

Is Qwen2.5 Plus 1127 better than Wizardlm 13b?

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 31.4 on the Noometry Index.

Is Qwen2.5 Plus 1127 or Wizardlm 13b better for coding?

Qwen2.5 Plus 1127 scores higher on coding benchmarks: 38.5 versus 30.1 in the Noometry coding category.

How many benchmarks do Qwen2.5 Plus 1127 and Wizardlm 13b share?

10 benchmarks have published results for both models. Qwen2.5 Plus 1127 has 14 scored results on Noometry and Wizardlm 13b has 10.

Related comparisons

Go deeper