Model comparison

Qwen3-4B vs Wizardlm 70b

Wizardlm 70b is the stronger model overall, scoring 33.0 to 31.9 on the Noometry Index.

Last verified . 0 shared benchmarks.

Qwen3-4B Alibaba (Qwen)

31.9

Rank #264 Confirmed

Wizardlm 70b Microsoft

33.0

Rank #249 Confirmed

Side by side

Qwen3-4B and Wizardlm 70b specifications
Qwen3-4BWizardlm 70b
ProviderAlibaba (Qwen)Microsoft
Noometry Index31.933.0
Released2025-04-29—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked612

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Qwen3-4B: —, Wizardlm 70b: 31.4 (#285)

Coding benchmarks
BenchmarkQwen3-4BWizardlm 70b
LMArena Coding—1081

Agentic & Tool Use Not comparable

Qwen3-4B: 27.6 (#100), Wizardlm 70b: —

Agentic & Tool Use benchmarks
BenchmarkQwen3-4BWizardlm 70b
Berkeley Function Calling Leaderboard35.7%—

Reasoning Wizardlm 70b leads

Qwen3-4B: 19.2 (#268), Wizardlm 70b: 20.7 (#234)

Reasoning benchmarks
BenchmarkQwen3-4BWizardlm 70b
Chess Puzzles4%—
LMArena Hard Prompts—1079

Math Wizardlm 70b leads

Qwen3-4B: 29.7 (#240), Wizardlm 70b: 32.2 (#218)

Math benchmarks
BenchmarkQwen3-4BWizardlm 70b
MathArena Final-Answer Competitions38.5%—
OTIS Mock AIME 2024-202552.2%—
LMArena Math—1116

Knowledge Not comparable

Qwen3-4B: 33.0 (#208), Wizardlm 70b: —

Knowledge benchmarks
BenchmarkQwen3-4BWizardlm 70b
GPQA Diamond52.3%—
Vectara Hallucination Rate5.7%—

Multilingual Not comparable

Qwen3-4B: —, Wizardlm 70b: 29.6 (#265)

Multilingual benchmarks
BenchmarkQwen3-4BWizardlm 70b
LMArena Non-English—1078
LMArena Chinese—1052
LMArena German—1083
LMArena Russian—1155

Instruction Following Not comparable

Qwen3-4B: —, Wizardlm 70b: 56.3 (#273)

Instruction Following benchmarks
BenchmarkQwen3-4BWizardlm 70b
LMArena Instruction Following—1093

Long Context Not comparable

Qwen3-4B: —, Wizardlm 70b: 33.3 (#263)

Long Context benchmarks
BenchmarkQwen3-4BWizardlm 70b
LMArena Longer Query—1097

Writing & Preference Not comparable

Qwen3-4B: —, Wizardlm 70b: 34.8 (#269)

Writing & Preference benchmarks
BenchmarkQwen3-4BWizardlm 70b
LMArena Text—1120
LMArena Creative Writing—1149
LMArena Multi-Turn—1108

Frequently asked questions

Is Qwen3-4B better than Wizardlm 70b?

Wizardlm 70b is the stronger model overall, scoring 33.0 to 31.9 on the Noometry Index.

How many benchmarks do Qwen3-4B and Wizardlm 70b share?

0 benchmarks have published results for both models. Qwen3-4B has 6 scored results on Noometry and Wizardlm 70b has 12.

Related comparisons

Go deeper