Model comparison

Qwen Max vs Wizardlm 70b

Qwen Max is the stronger model overall, scoring 34.7 to 33.0 on the Noometry Index.

Last verified . 12 shared benchmarks.

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Wizardlm 70b Microsoft

33.0

Rank #249 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Qwen Max scores higher in 5 categories and Wizardlm 70b in 2 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen Max leads 47.8 to 34.8.
  • Wizardlm 70b has downloadable open weights; the other is API-only.

Side by side

Qwen Max and Wizardlm 70b specifications
Qwen MaxWizardlm 70b
ProviderAlibaba (Qwen)Microsoft
Noometry Index34.733.0
Released2024-04-03—
WeightsProprietaryOpen
Context window33K—
Max output8K—
Input $ / M tokens$1.60—
Output $ / M tokens$6.40—
Results tracked2312

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Qwen Max: 30.7 (#292), Wizardlm 70b: 31.4 (#285)

Coding benchmarks
BenchmarkQwen MaxWizardlm 70b
LMArena Coding12881081
Aider Polyglot21.8%—

Reasoning Qwen Max leads

Qwen Max: 25.1 (#151), Wizardlm 70b: 20.7 (#234)

Reasoning benchmarks
BenchmarkQwen MaxWizardlm 70b
LMArena Hard Prompts12691079

Math Wizardlm 70b leads

Qwen Max: 22.3 (#276), Wizardlm 70b: 32.2 (#218)

Math benchmarks
BenchmarkQwen MaxWizardlm 70b
LMArena Math12751116
OTIS Mock AIME 2024-202516.1%—
MATH Level 567.2%—
FrontierMath (Feb 2025 set)1%—

Knowledge Not comparable

Qwen Max: 30.3 (#228), Wizardlm 70b: —

Knowledge benchmarks
BenchmarkQwen MaxWizardlm 70b
GPQA Diamond56.1%—
LMArena Expert1248—

Multilingual Qwen Max leads

Qwen Max: 41.8 (#202), Wizardlm 70b: 29.6 (#265)

Multilingual benchmarks
BenchmarkQwen MaxWizardlm 70b
LMArena Non-English12631078
LMArena Chinese12541052
LMArena German12541083
LMArena Russian12741155
LMArena French1330—
LMArena Japanese1205—
LMArena Korean1142—
LMArena Spanish1290—

Instruction Following Qwen Max leads

Qwen Max: 66.5 (#208), Wizardlm 70b: 56.3 (#273)

Instruction Following benchmarks
BenchmarkQwen MaxWizardlm 70b
LMArena Instruction Following12621093

Long Context Qwen Max leads

Qwen Max: 39.4 (#180), Wizardlm 70b: 33.3 (#263)

Long Context benchmarks
BenchmarkQwen MaxWizardlm 70b
LMArena Longer Query12881097
Fiction.LiveBench66.7%—

Writing & Preference Qwen Max leads

Qwen Max: 47.8 (#205), Wizardlm 70b: 34.8 (#269)

Writing & Preference benchmarks
BenchmarkQwen MaxWizardlm 70b
LMArena Text12821120
LMArena Creative Writing12481149
LMArena Multi-Turn12771108

Frequently asked questions

Is Qwen Max better than Wizardlm 70b?

Qwen Max is the stronger model overall, scoring 34.7 to 33.0 on the Noometry Index.

Is Qwen Max or Wizardlm 70b better for coding?

They score almost the same on coding (30.7 vs 31.4); test both on your own repository before choosing.

How many benchmarks do Qwen Max and Wizardlm 70b share?

12 benchmarks have published results for both models. Qwen Max has 23 scored results on Noometry and Wizardlm 70b has 12.

Related comparisons

Go deeper