Model comparison

Qwen3.5-9B vs Wizardlm 13b

Qwen3.5-9B is the stronger model overall, scoring 33.8 to 31.4 on the Noometry Index.

Last verified . 0 shared benchmarks.

Qwen3.5-9B Alibaba (Qwen)

33.8

Rank #236 Confirmed

Wizardlm 13b Microsoft

31.4

Rank #274 Confirmed

Summary

  • The widest gap is in coding, where Qwen3.5-9B leads 35.9 to 30.1.

Side by side

Qwen3.5-9B and Wizardlm 13b specifications
Qwen3.5-9BWizardlm 13b
ProviderAlibaba (Qwen)Microsoft
Noometry Index33.831.4
Released2026-02-23—
WeightsOpenOpen
Context window262K—
Max output66K—
Input $ / M tokens$0.10—
Output $ / M tokens$0.15—
Results tracked1010

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.5-9B leads

Qwen3.5-9B: 35.9 (#217), Wizardlm 13b: 30.1 (#298)

Coding benchmarks
BenchmarkQwen3.5-9BWizardlm 13b
SciCode27.5%—
LMArena Coding—1035

Agentic & Tool Use Not comparable

Qwen3.5-9B: 14.5 (#151), Wizardlm 13b: —

Agentic & Tool Use benchmarks
BenchmarkQwen3.5-9BWizardlm 13b
Terminal-Bench9.2%—

Reasoning Qwen3.5-9B leads

Qwen3.5-9B: 23.1 (#182), Wizardlm 13b: 19.4 (#259)

Reasoning benchmarks
BenchmarkQwen3.5-9BWizardlm 13b
CritPt0.3%—
Chess Puzzles12%—
LMArena Hard Prompts—1018
DTBench71.2%—
LMCA24.5%—
Epoch Capabilities Index139.46—

Math Qwen3.5-9B leads

Qwen3.5-9B: 34.8 (#192), Wizardlm 13b: 30.2 (#238)

Math benchmarks
BenchmarkQwen3.5-9BWizardlm 13b
MathArena Final-Answer Competitions48.5%—
OTIS Mock AIME 2024-202561.7%—
LMArena Math—1017

Knowledge Not comparable

Qwen3.5-9B: 46.0 (#84), Wizardlm 13b: —

Knowledge benchmarks
BenchmarkQwen3.5-9BWizardlm 13b
GPQA Diamond79%—

Multilingual Not comparable

Qwen3.5-9B: —, Wizardlm 13b: 27.1 (#277)

Multilingual benchmarks
BenchmarkQwen3.5-9BWizardlm 13b
LMArena Non-English—1034
LMArena Chinese—1023

Instruction Following Not comparable

Qwen3.5-9B: —, Wizardlm 13b: 53.5 (#285)

Instruction Following benchmarks
BenchmarkQwen3.5-9BWizardlm 13b
LMArena Instruction Following—1048

Long Context Not comparable

Qwen3.5-9B: —, Wizardlm 13b: 32.0 (#273)

Long Context benchmarks
BenchmarkQwen3.5-9BWizardlm 13b
LMArena Longer Query—1054

Writing & Preference Not comparable

Qwen3.5-9B: —, Wizardlm 13b: 30.5 (#287)

Writing & Preference benchmarks
BenchmarkQwen3.5-9BWizardlm 13b
LMArena Text—1077
LMArena Creative Writing—1091
LMArena Multi-Turn—1047

Frequently asked questions

Is Qwen3.5-9B better than Wizardlm 13b?

Qwen3.5-9B is the stronger model overall, scoring 33.8 to 31.4 on the Noometry Index.

Is Qwen3.5-9B or Wizardlm 13b better for coding?

Qwen3.5-9B scores higher on coding benchmarks: 35.9 versus 30.1 in the Noometry coding category.

How many benchmarks do Qwen3.5-9B and Wizardlm 13b share?

0 benchmarks have published results for both models. Qwen3.5-9B has 10 scored results on Noometry and Wizardlm 13b has 10.

Related comparisons

Go deeper