Model comparison

Qwen Plus vs Qwen3-4B

Qwen Plus is the stronger model overall, scoring 37.1 to 31.9 on the Noometry Index.

Last verified . 2 shared benchmarks.

Qwen Plus Alibaba (Qwen)

37.1

Rank #210 Confirmed

Qwen3-4B Alibaba (Qwen)

31.9

Rank #264 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Qwen Plus scores higher in 1 category and Qwen3-4B in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen Plus leads 28.4 to 19.2.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 17.8% for Qwen Plus and 52.2% for Qwen3-4B.
  • Qwen3-4B has downloadable open weights; the other is API-only.

Side by side

Qwen Plus and Qwen3-4B specifications
Qwen PlusQwen3-4B
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index37.131.9
Released2024-01-252025-04-29
WeightsProprietaryOpen
Context window1M—
Max output33K—
Input $ / M tokens$0.40—
Output $ / M tokens$1.20—
Results tracked206

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Qwen Plus: 38.9 (#167), Qwen3-4B: —

Coding benchmarks
BenchmarkQwen PlusQwen3-4B
LMArena Coding1328—

Agentic & Tool Use Not comparable

Qwen Plus: —, Qwen3-4B: 27.6 (#100)

Agentic & Tool Use benchmarks
BenchmarkQwen PlusQwen3-4B
Berkeley Function Calling Leaderboard—35.7%

Reasoning Qwen Plus leads

Qwen Plus: 28.4 (#107), Qwen3-4B: 19.2 (#268)

Reasoning benchmarks
BenchmarkQwen PlusQwen3-4B
Kagi LLM Benchmark63.3%—
Chess Puzzles—4%
LMArena Hard Prompts1317—
DTBench81.1%—
LMCA24%—

Math Qwen3-4B leads

Qwen Plus: 23.3 (#271), Qwen3-4B: 29.7 (#240)

Math benchmarks
BenchmarkQwen PlusQwen3-4B
OTIS Mock AIME 2024-202517.8%52.2%
MathArena Final-Answer Competitions—38.5%
LMArena Math1326—
MATH Level 565.3%—
FrontierMath (Feb 2025 set)1.7%—

Knowledge Qwen3-4B leads

Qwen Plus: 27.4 (#251), Qwen3-4B: 33.0 (#208)

Knowledge benchmarks
BenchmarkQwen PlusQwen3-4B
GPQA Diamond48.1%52.3%
Vectara Hallucination Rate—5.7%
LMArena Expert1328—

Multilingual Not comparable

Qwen Plus: 45.1 (#175), Qwen3-4B: —

Multilingual benchmarks
BenchmarkQwen PlusQwen3-4B
LMArena Non-English1310—
LMArena Chinese1347—
LMArena Japanese1251—
LMArena Russian1323—

Instruction Following Not comparable

Qwen Plus: 68.8 (#181), Qwen3-4B: —

Instruction Following benchmarks
BenchmarkQwen PlusQwen3-4B
LMArena Instruction Following1303—

Long Context Not comparable

Qwen Plus: 40.3 (#158), Qwen3-4B: —

Long Context benchmarks
BenchmarkQwen PlusQwen3-4B
LMArena Longer Query1324—

Writing & Preference Not comparable

Qwen Plus: 52.2 (#176), Qwen3-4B: —

Writing & Preference benchmarks
BenchmarkQwen PlusQwen3-4B
LMArena Text1326—
LMArena Creative Writing1293—
LMArena Multi-Turn1336—

Frequently asked questions

Is Qwen Plus better than Qwen3-4B?

Qwen Plus is the stronger model overall, scoring 37.1 to 31.9 on the Noometry Index.

How many benchmarks do Qwen Plus and Qwen3-4B share?

2 benchmarks have published results for both models. Qwen Plus has 20 scored results on Noometry and Qwen3-4B has 6.

Related comparisons

Go deeper