Model comparison

Qwen3-30B-A3B vs Qwen3.5-9B

Qwen3-30B-A3B is the stronger model overall, scoring 38.9 to 33.8 on the Noometry Index. Qwen3.5-9B costs 1.9× less per token, which makes it the better buy when Qwen3-30B-A3B's lead doesn't matter for your workload.

Last verified . 9 shared benchmarks.

Qwen3-30B-A3B Alibaba (Qwen)

38.9

Rank #179 Confirmed

Qwen3.5-9B Alibaba (Qwen)

33.8

Rank #236 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Qwen3-30B-A3B scores higher in 3 categories and Qwen3.5-9B in 2 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where Qwen3-30B-A3B leads 29.8 to 14.5.
  • The biggest single-benchmark swing is GPQA Diamond: 70.1% for Qwen3-30B-A3B and 79% for Qwen3.5-9B.
  • Qwen3.5-9B is cheaper at $0.10 / $0.15 per million input/output tokens, against $0.12 / $0.50 for Qwen3-30B-A3B.
  • Qwen3.5-9B accepts more context: 262K tokens versus 41K.

Side by side

Qwen3-30B-A3B and Qwen3.5-9B specifications
Qwen3-30B-A3BQwen3.5-9B
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index38.933.8
Released2025-04-282026-02-23
WeightsOpenOpen
Context window41K262K
Max output16K66K
Input $ / M tokens$0.12$0.10
Output $ / M tokens$0.50$0.15
Results tracked3210

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3-30B-A3B leads

Qwen3-30B-A3B: 37.5 (#194), Qwen3.5-9B: 35.9 (#217)

Coding benchmarks
BenchmarkQwen3-30B-A3BQwen3.5-9B
SciCode33.3%27.5%
WeirdML29.8%—
LMArena Coding1416—

Agentic & Tool Use Qwen3-30B-A3B leads

Qwen3-30B-A3B: 29.8 (#82), Qwen3.5-9B: 14.5 (#151)

Agentic & Tool Use benchmarks
BenchmarkQwen3-30B-A3BQwen3.5-9B
Terminal-Bench—9.2%
Berkeley Function Calling Leaderboard41.4%—

Reasoning Too close to call

Qwen3-30B-A3B: 22.2 (#204), Qwen3.5-9B: 23.1 (#182)

Reasoning benchmarks
BenchmarkQwen3-30B-A3BQwen3.5-9B
CritPt0.3%0.3%
Chess Puzzles8%12%
DTBench69.3%71.2%
LMCA22.4%24.5%
Epoch Capabilities Index139.63139.46
Kagi LLM Benchmark54.9%—
LMArena Hard Prompts1398—

Math Qwen3-30B-A3B leads

Qwen3-30B-A3B: 37.4 (#157), Qwen3.5-9B: 34.8 (#192)

Math benchmarks
BenchmarkQwen3-30B-A3BQwen3.5-9B
MathArena Final-Answer Competitions47.8%48.5%
OTIS Mock AIME 2024-202570.3%61.7%
LMArena Math1394—

Knowledge Qwen3.5-9B leads

Qwen3-30B-A3B: 41.8 (#105), Qwen3.5-9B: 46.0 (#84)

Knowledge benchmarks
BenchmarkQwen3-30B-A3BQwen3.5-9B
GPQA Diamond70.1%79%
Confabulations12.3%—
LMArena Expert1396—

Multilingual Not comparable

Qwen3-30B-A3B: 49.5 (#132), Qwen3.5-9B: —

Multilingual benchmarks
BenchmarkQwen3-30B-A3BQwen3.5-9B
LMArena Non-English1372—
LMArena Chinese1433—
LMArena French1418—
LMArena German1380—
LMArena Japanese1337—
LMArena Korean1331—
LMArena Russian1370—
LMArena Spanish1404—

Instruction Following Not comparable

Qwen3-30B-A3B: 72.0 (#142), Qwen3.5-9B: —

Instruction Following benchmarks
BenchmarkQwen3-30B-A3BQwen3.5-9B
LMArena Instruction Following1363—

Long Context Not comparable

Qwen3-30B-A3B: 31.0 (#283), Qwen3.5-9B: —

Long Context benchmarks
BenchmarkQwen3-30B-A3BQwen3.5-9B
Fiction.LiveBench40.6%—
LMArena Longer Query1379—

Writing & Preference Not comparable

Qwen3-30B-A3B: 55.6 (#143), Qwen3.5-9B: —

Writing & Preference benchmarks
BenchmarkQwen3-30B-A3BQwen3.5-9B
LMArena Text1384—
LMArena Creative Writing1317—
Short-Story Creative Writing75.3%—
LMArena Multi-Turn1378—

Frequently asked questions

Is Qwen3-30B-A3B better than Qwen3.5-9B?

Qwen3-30B-A3B is the stronger model overall, scoring 38.9 to 33.8 on the Noometry Index. Qwen3.5-9B costs 1.9× less per token, which makes it the better buy when Qwen3-30B-A3B's lead doesn't matter for your workload.

Which is cheaper, Qwen3-30B-A3B or Qwen3.5-9B?

Qwen3.5-9B is cheaper. It lists at $0.10 per million input tokens and $0.15 per million output tokens; Qwen3-30B-A3B lists at $0.12 and $0.50.

Is Qwen3-30B-A3B or Qwen3.5-9B better for coding?

Qwen3-30B-A3B scores higher on coding benchmarks: 37.5 versus 35.9 in the Noometry coding category.

Which has the bigger context window?

Qwen3.5-9B does, with 262K tokens against 41K.

How many benchmarks do Qwen3-30B-A3B and Qwen3.5-9B share?

9 benchmarks have published results for both models. Qwen3-30B-A3B has 32 scored results on Noometry and Qwen3.5-9B has 10.

Related comparisons

Go deeper