Model comparison

Qwen3.5-9B vs Wizardlm 70b

Qwen3.5-9B and Wizardlm 70b score almost the same on the Noometry Index (33.8 vs 33.0), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

Qwen3.5-9B Alibaba (Qwen)

33.8

Rank #236 Confirmed

Wizardlm 70b Microsoft

33.0

Rank #249 Confirmed

Summary

  • The widest gap is in coding, where Qwen3.5-9B leads 35.9 to 31.4.

Side by side

Qwen3.5-9B and Wizardlm 70b specifications
Qwen3.5-9BWizardlm 70b
ProviderAlibaba (Qwen)Microsoft
Noometry Index33.833.0
Released2026-02-23—
WeightsOpenOpen
Context window262K—
Max output66K—
Input $ / M tokens$0.10—
Output $ / M tokens$0.15—
Results tracked1012

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.5-9B leads

Qwen3.5-9B: 35.9 (#217), Wizardlm 70b: 31.4 (#285)

Coding benchmarks
BenchmarkQwen3.5-9BWizardlm 70b
SciCode27.5%—
LMArena Coding—1081

Agentic & Tool Use Not comparable

Qwen3.5-9B: 14.5 (#151), Wizardlm 70b: —

Agentic & Tool Use benchmarks
BenchmarkQwen3.5-9BWizardlm 70b
Terminal-Bench9.2%—

Reasoning Qwen3.5-9B leads

Qwen3.5-9B: 23.1 (#182), Wizardlm 70b: 20.7 (#234)

Reasoning benchmarks
BenchmarkQwen3.5-9BWizardlm 70b
CritPt0.3%—
Chess Puzzles12%—
LMArena Hard Prompts—1079
DTBench71.2%—
LMCA24.5%—
Epoch Capabilities Index139.46—

Math Qwen3.5-9B leads

Qwen3.5-9B: 34.8 (#192), Wizardlm 70b: 32.2 (#218)

Math benchmarks
BenchmarkQwen3.5-9BWizardlm 70b
MathArena Final-Answer Competitions48.5%—
OTIS Mock AIME 2024-202561.7%—
LMArena Math—1116

Knowledge Not comparable

Qwen3.5-9B: 46.0 (#84), Wizardlm 70b: —

Knowledge benchmarks
BenchmarkQwen3.5-9BWizardlm 70b
GPQA Diamond79%—

Multilingual Not comparable

Qwen3.5-9B: —, Wizardlm 70b: 29.6 (#265)

Multilingual benchmarks
BenchmarkQwen3.5-9BWizardlm 70b
LMArena Non-English—1078
LMArena Chinese—1052
LMArena German—1083
LMArena Russian—1155

Instruction Following Not comparable

Qwen3.5-9B: —, Wizardlm 70b: 56.3 (#273)

Instruction Following benchmarks
BenchmarkQwen3.5-9BWizardlm 70b
LMArena Instruction Following—1093

Long Context Not comparable

Qwen3.5-9B: —, Wizardlm 70b: 33.3 (#263)

Long Context benchmarks
BenchmarkQwen3.5-9BWizardlm 70b
LMArena Longer Query—1097

Writing & Preference Not comparable

Qwen3.5-9B: —, Wizardlm 70b: 34.8 (#269)

Writing & Preference benchmarks
BenchmarkQwen3.5-9BWizardlm 70b
LMArena Text—1120
LMArena Creative Writing—1149
LMArena Multi-Turn—1108

Frequently asked questions

Is Qwen3.5-9B better than Wizardlm 70b?

Qwen3.5-9B and Wizardlm 70b score almost the same on the Noometry Index (33.8 vs 33.0), so choose on price, context window or the category you care about most.

Is Qwen3.5-9B or Wizardlm 70b better for coding?

Qwen3.5-9B scores higher on coding benchmarks: 35.9 versus 31.4 in the Noometry coding category.

How many benchmarks do Qwen3.5-9B and Wizardlm 70b share?

0 benchmarks have published results for both models. Qwen3.5-9B has 10 scored results on Noometry and Wizardlm 70b has 12.

Related comparisons

Go deeper