Model comparison

Muse Spark 1.1 vs Qwen3-30B-A3B

Muse Spark 1.1 is the stronger model overall, scoring 49.9 to 38.9 on the Noometry Index. Qwen3-30B-A3B costs 9.3× less per token, which makes it the better buy when Muse Spark 1.1's lead doesn't matter for your workload.

Last verified . 22 shared benchmarks.

Muse Spark 1.1 Meta

49.9

Rank #51 Confirmed

Qwen3-30B-A3B Alibaba (Qwen)

38.9

Rank #179 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Muse Spark 1.1 scores higher in 9 categories and Qwen3-30B-A3B in 0 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Muse Spark 1.1 leads 47.1 to 22.2.
  • The biggest single-benchmark swing is LMCA: 49.9% for Muse Spark 1.1 and 22.4% for Qwen3-30B-A3B.
  • Qwen3-30B-A3B is cheaper at $0.12 / $0.50 per million input/output tokens, against $1.25 / $4.25 for Muse Spark 1.1.
  • Muse Spark 1.1 accepts more context: 1.05M tokens versus 41K.
  • Qwen3-30B-A3B has downloadable open weights; the other is API-only.

Side by side

Muse Spark 1.1 and Qwen3-30B-A3B specifications
Muse Spark 1.1Qwen3-30B-A3B
ProviderMetaAlibaba (Qwen)
Noometry Index49.938.9
Released2026-04-082025-04-28
WeightsProprietaryOpen
Context window1.05M41K
Max output131K16K
Input $ / M tokens$1.25$0.12
Output $ / M tokens$4.25$0.50
Results tracked3732

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Muse Spark 1.1 leads

Muse Spark 1.1: 51.3 (#40), Qwen3-30B-A3B: 37.5 (#194)

Coding benchmarks
BenchmarkMuse Spark 1.1Qwen3-30B-A3B
SciCode58.8%33.3%
LMArena Coding14981416
DeepSWE53.3%—
LMArena WebDev1542—
WeirdML—29.8%

Agentic & Tool Use Muse Spark 1.1 leads

Muse Spark 1.1: 30.8 (#73), Qwen3-30B-A3B: 29.8 (#82)

Agentic & Tool Use benchmarks
BenchmarkMuse Spark 1.1Qwen3-30B-A3B
APEX-Agents31.8%—
Berkeley Function Calling Leaderboard—41.4%
τ²-bench Banking40.5%—
GBAEval7.9%—
GDP.pdf15%—
Vending-Bench 26,520—

Reasoning Muse Spark 1.1 leads

Muse Spark 1.1: 47.1 (#44), Qwen3-30B-A3B: 22.2 (#204)

Reasoning benchmarks
BenchmarkMuse Spark 1.1Qwen3-30B-A3B
CritPt15.1%0.3%
LMArena Hard Prompts14861398
DTBench94.4%69.3%
LMCA49.9%22.4%
Epoch Capabilities Index154.21139.63
Kagi LLM Benchmark—54.9%
NYT Connections (extended)84.9%—
Chess Puzzles—8%
Surface Evolver Bench52.5%—

Math Muse Spark 1.1 leads

Muse Spark 1.1: 45.5 (#76), Qwen3-30B-A3B: 37.4 (#157)

Math benchmarks
BenchmarkMuse Spark 1.1Qwen3-30B-A3B
LMArena Math14831394
MathArena Final-Answer Competitions—47.8%
OTIS Mock AIME 2024-2025—70.3%
ProofBench39%—

Knowledge Muse Spark 1.1 leads

Muse Spark 1.1: 53.1 (#59), Qwen3-30B-A3B: 41.8 (#105)

Knowledge benchmarks
BenchmarkMuse Spark 1.1Qwen3-30B-A3B
LMArena Expert14781396
GPQA Diamond—70.1%
SimpleQA Verified57.8%—
Confabulations—12.3%

Multimodal Not comparable

Muse Spark 1.1: 42.6 (#29), Qwen3-30B-A3B: —

Multimodal benchmarks
BenchmarkMuse Spark 1.1Qwen3-30B-A3B
LMArena Vision1293—
LMArena Document1465—

Multilingual Muse Spark 1.1 leads

Muse Spark 1.1: 56.7 (#17), Qwen3-30B-A3B: 49.5 (#132)

Multilingual benchmarks
BenchmarkMuse Spark 1.1Qwen3-30B-A3B
LMArena Non-English14721372
LMArena Chinese15181433
LMArena French14941418
LMArena German14661380
LMArena Japanese14511337
LMArena Korean14581331
LMArena Russian14831370
LMArena Spanish14641404

Instruction Following Muse Spark 1.1 leads

Muse Spark 1.1: 76.5 (#39), Qwen3-30B-A3B: 72.0 (#142)

Instruction Following benchmarks
BenchmarkMuse Spark 1.1Qwen3-30B-A3B
LMArena Instruction Following14571363

Long Context Muse Spark 1.1 leads

Muse Spark 1.1: 44.8 (#58), Qwen3-30B-A3B: 31.0 (#283)

Long Context benchmarks
BenchmarkMuse Spark 1.1Qwen3-30B-A3B
LMArena Longer Query14621379
Fiction.LiveBench—40.6%

Writing & Preference Muse Spark 1.1 leads

Muse Spark 1.1: 73.4 (#11), Qwen3-30B-A3B: 55.6 (#143)

Writing & Preference benchmarks
BenchmarkMuse Spark 1.1Qwen3-30B-A3B
LMArena Text14791384
LMArena Creative Writing14371317
LMArena Multi-Turn14851378
Short-Story Creative Writing—75.3%
EQ-Bench Creative Writing1927—
EQ-Bench 41260—

Frequently asked questions

Is Muse Spark 1.1 better than Qwen3-30B-A3B?

Muse Spark 1.1 is the stronger model overall, scoring 49.9 to 38.9 on the Noometry Index. Qwen3-30B-A3B costs 9.3× less per token, which makes it the better buy when Muse Spark 1.1's lead doesn't matter for your workload.

Which is cheaper, Muse Spark 1.1 or Qwen3-30B-A3B?

Qwen3-30B-A3B is cheaper. It lists at $0.12 per million input tokens and $0.50 per million output tokens; Muse Spark 1.1 lists at $1.25 and $4.25.

Is Muse Spark 1.1 or Qwen3-30B-A3B better for coding?

Muse Spark 1.1 scores higher on coding benchmarks: 51.3 versus 37.5 in the Noometry coding category.

Which has the bigger context window?

Muse Spark 1.1 does, with 1.05M tokens against 41K.

How many benchmarks do Muse Spark 1.1 and Qwen3-30B-A3B share?

22 benchmarks have published results for both models. Muse Spark 1.1 has 37 scored results on Noometry and Qwen3-30B-A3B has 32.

Related comparisons

Go deeper