Alibaba (Qwen), open weights

Qwen3-4B

Qwen3-4B by Alibaba (Qwen) ranks 264th of 354 ranked models on the Noometry Index as of October 2026, with a score of 31.9. Its strongest category is agentic & tool use, where it ranks 100th.

Last verified

Specifications

Noometry rank
#264 of 354
Index score
31.9
Evidence
Confirmed 6 results
Released
April 29, 2025
Weights
Open weights
Reasoning
Unknown
Context window
—
Max output
—
Input price
Not listed
Output price
Not listed
Blended price
Not listed
Output speed
Not measured
Value
Not ranked
Knowledge cutoff
Unknown

Category scores

Each category score combines every public result we have in that category.

Qwen3-4B category scores
  1. Agentic & Tool Use 27.6
  2. Reasoning 19.2
  3. Math 29.7
  4. Knowledge 33.0
Qwen3-4B category ranks
CategoryScoreRankResults
Agentic & Tool Use27.6#1001
Reasoning19.2#2681
Math29.7#2402
Knowledge33.0#2082

Strengths and weaknesses

Categories where Qwen3-4B places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Qwen3-4B: strongest categories
CategoryScorevs medianRank
Agentic & Tool Use27.6−2.8#100 of 154, top 65%
Knowledge33.0−4.3#208 of 314, top 67%

Weakest categories

Qwen3-4B: weakest categories
CategoryScorevs medianRank
Reasoning19.2−4.4#268 of 350, top 77%
Math29.7−6.9#240 of 327, top 74%

Closest competitors

The models ranked just above and below Qwen3-4B. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Qwen3-4B
ModelRankScoreBlended $/MSpeed
Falcon-180B#26032.2——Compare
Gemini 1.5 Pro (May 2024)#26132.1——Compare
Gemma 3 12B#26232.1$0.075—Compare
Mistral Large#26331.9$3—Compare
Amazon Nova Lite#26531.9$0.10—Compare
Mistral Medium 3.1#26631.9$0.80—Compare
Qwen2.5 72B Instruct#26731.9$2.45—Compare
Llama2 70b Steerlm Chat#26831.8——Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Agentic & Tool Use

Qwen3-4B Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Berkeley Function Calling Leaderboard35.7%#29 of 49, top 60%fcBerkeley Function Calling Leaderboard

Reasoning

Qwen3-4B Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
Chess Puzzles4%#100 of 129, top 78%Epoch AI2026-08-28
Chess Puzzles1%noneEpoch AI2026-08-28

Math

Qwen3-4B Math benchmark results
BenchmarkScorePositionSettingSourceDate
MathArena Final-Answer Competitions38.5%#29 of 29, top 100%MathArena
OTIS Mock AIME 2024-202552.2%#111 of 173, top 65%Epoch AI2026-08-28

Knowledge

Qwen3-4B Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond52.3%#124 of 186, top 67%Epoch AI2026-08-28
GPQA Diamond45.8%Epoch AI2026-08-28
GPQA Diamond43.6%noneEpoch AI2026-08-28
Vectara Hallucination Rate (lower is better)5.7%#19 of 96, top 20%Vectara Hallucination Leaderboard

Compare Qwen3-4B

Other Alibaba (Qwen) models

Frequently asked questions

How good is Qwen3-4B?

Qwen3-4B by Alibaba (Qwen) ranks 264th of 354 ranked models on the Noometry Index as of October 2026, with a score of 31.9. Its strongest category is agentic & tool use, where it ranks 100th.

Is Qwen3-4B open source?

Yes. Qwen3-4B's weights are downloadable; check the license for commercial terms.

What are Qwen3-4B's strengths and weaknesses?

Relative to other ranked models, Qwen3-4B places best in agentic & tool use, knowledge and lowest in reasoning, math.

What is Qwen3-4B best at?

Its best category is agentic & tool use, where it ranks 100th on Noometry.