Model comparison

Qwen2.5 Plus 1127 vs Wizardlm 70b

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 33.0 on the Noometry Index.

Last verified . 12 shared benchmarks.

Qwen2.5 Plus 1127 Alibaba (Qwen)

38.8

Rank #181 Confirmed

Wizardlm 70b Microsoft

33.0

Rank #249 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Qwen2.5 Plus 1127 scores higher in 7 categories and Wizardlm 70b in 0 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen2.5 Plus 1127 leads 49.4 to 34.8.
  • Wizardlm 70b has downloadable open weights; the other is API-only.

Side by side

Qwen2.5 Plus 1127 and Wizardlm 70b specifications
Qwen2.5 Plus 1127Wizardlm 70b
ProviderAlibaba (Qwen)Microsoft
Noometry Index38.833.0
Released——
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1412

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 Plus 1127 leads

Qwen2.5 Plus 1127: 38.5 (#175), Wizardlm 70b: 31.4 (#285)

Coding benchmarks
BenchmarkQwen2.5 Plus 1127Wizardlm 70b
LMArena Coding13141081

Reasoning Qwen2.5 Plus 1127 leads

Qwen2.5 Plus 1127: 25.9 (#141), Wizardlm 70b: 20.7 (#234)

Reasoning benchmarks
BenchmarkQwen2.5 Plus 1127Wizardlm 70b
LMArena Hard Prompts12991079

Math Qwen2.5 Plus 1127 leads

Qwen2.5 Plus 1127: 36.1 (#174), Wizardlm 70b: 32.2 (#218)

Math benchmarks
BenchmarkQwen2.5 Plus 1127Wizardlm 70b
LMArena Math12981116

Knowledge Not comparable

Qwen2.5 Plus 1127: 35.5 (#183), Wizardlm 70b: —

Knowledge benchmarks
BenchmarkQwen2.5 Plus 1127Wizardlm 70b
LMArena Expert1289—

Multilingual Qwen2.5 Plus 1127 leads

Qwen2.5 Plus 1127: 41.9 (#201), Wizardlm 70b: 29.6 (#265)

Multilingual benchmarks
BenchmarkQwen2.5 Plus 1127Wizardlm 70b
LMArena Non-English12651078
LMArena Chinese13141052
LMArena German12311083
LMArena Russian12711155
LMArena Japanese1207—

Instruction Following Qwen2.5 Plus 1127 leads

Qwen2.5 Plus 1127: 67.2 (#199), Wizardlm 70b: 56.3 (#273)

Instruction Following benchmarks
BenchmarkQwen2.5 Plus 1127Wizardlm 70b
LMArena Instruction Following12751093

Long Context Qwen2.5 Plus 1127 leads

Qwen2.5 Plus 1127: 39.2 (#184), Wizardlm 70b: 33.3 (#263)

Long Context benchmarks
BenchmarkQwen2.5 Plus 1127Wizardlm 70b
LMArena Longer Query12921097

Writing & Preference Qwen2.5 Plus 1127 leads

Qwen2.5 Plus 1127: 49.4 (#192), Wizardlm 70b: 34.8 (#269)

Writing & Preference benchmarks
BenchmarkQwen2.5 Plus 1127Wizardlm 70b
LMArena Text12991120
LMArena Creative Writing12621149
LMArena Multi-Turn12991108

Frequently asked questions

Is Qwen2.5 Plus 1127 better than Wizardlm 70b?

Qwen2.5 Plus 1127 is the stronger model overall, scoring 38.8 to 33.0 on the Noometry Index.

Is Qwen2.5 Plus 1127 or Wizardlm 70b better for coding?

Qwen2.5 Plus 1127 scores higher on coding benchmarks: 38.5 versus 31.4 in the Noometry coding category.

How many benchmarks do Qwen2.5 Plus 1127 and Wizardlm 70b share?

12 benchmarks have published results for both models. Qwen2.5 Plus 1127 has 14 scored results on Noometry and Wizardlm 70b has 12.

Related comparisons

Go deeper