Model comparison

Muse Spark vs Qwen3.5-9B

Muse Spark is the stronger model overall, scoring 50.6 to 33.8 on the Noometry Index.

Last verified . 5 shared benchmarks.

Muse Spark Meta

50.6

Rank #46 Confirmed

Qwen3.5-9B Alibaba (Qwen)

33.8

Rank #236 Confirmed

Summary

  • They share 5 benchmarks with published results for both. Muse Spark scores higher in 4 categories and Qwen3.5-9B in 0 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Muse Spark leads 65.7 to 46.0.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 88.9% for Muse Spark and 61.7% for Qwen3.5-9B.
  • Qwen3.5-9B has downloadable open weights; the other is API-only.

Side by side

Muse Spark and Qwen3.5-9B specifications
Muse SparkQwen3.5-9B
ProviderMetaAlibaba (Qwen)
Noometry Index50.633.8
Released2026-04-082026-02-23
WeightsProprietaryOpen
Context window—262K
Max output—66K
Input $ / M tokens—$0.10
Output $ / M tokens—$0.15
Results tracked2710

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Muse Spark leads

Muse Spark: 46.2 (#69), Qwen3.5-9B: 35.9 (#217)

Coding benchmarks
BenchmarkMuse SparkQwen3.5-9B
SciCode51.5%27.5%
LMArena Coding1481—

Agentic & Tool Use Not comparable

Muse Spark: —, Qwen3.5-9B: 14.5 (#151)

Agentic & Tool Use benchmarks
BenchmarkMuse SparkQwen3.5-9B
Terminal-Bench—9.2%

Reasoning Muse Spark leads

Muse Spark: 35.9 (#67), Qwen3.5-9B: 23.1 (#182)

Reasoning benchmarks
BenchmarkMuse SparkQwen3.5-9B
CritPt11.3%0.3%
Epoch Capabilities Index152.04139.46
Chess Puzzles—12%
LMArena Hard Prompts1474—
DTBench—71.2%
LMCA—24.5%

Math Muse Spark leads

Muse Spark: 47.8 (#66), Qwen3.5-9B: 34.8 (#192)

Math benchmarks
BenchmarkMuse SparkQwen3.5-9B
OTIS Mock AIME 2024-202588.9%61.7%
MathArena Final-Answer Competitions—48.5%
ProofBench17%—
LMArena Math1455—
FrontierMath (Feb 2025 set)39%—
FrontierMath Tier 4 (v1)14.6%—

Knowledge Muse Spark leads

Muse Spark: 65.7 (#13), Qwen3.5-9B: 46.0 (#84)

Knowledge benchmarks
BenchmarkMuse SparkQwen3.5-9B
GPQA Diamond89.8%79%
Humanity's Last Exam40.6%—
LMArena Expert1457—

Multimodal Not comparable

Muse Spark: 43.4 (#24), Qwen3.5-9B: —

Multimodal benchmarks
BenchmarkMuse SparkQwen3.5-9B
LMArena Vision1306—
LMArena Document1444—

Multilingual Not comparable

Muse Spark: 56.1 (#24), Qwen3.5-9B: —

Multilingual benchmarks
BenchmarkMuse SparkQwen3.5-9B
LMArena Non-English1464—
LMArena Chinese1509—
LMArena French1497—
LMArena German1497—
LMArena Korean1459—
LMArena Russian1466—
LMArena Spanish1472—

Instruction Following Not comparable

Muse Spark: 75.9 (#51), Qwen3.5-9B: —

Instruction Following benchmarks
BenchmarkMuse SparkQwen3.5-9B
LMArena Instruction Following1442—

Long Context Not comparable

Muse Spark: 44.4 (#69), Qwen3.5-9B: —

Long Context benchmarks
BenchmarkMuse SparkQwen3.5-9B
LMArena Longer Query1451—

Writing & Preference Not comparable

Muse Spark: 66.0 (#39), Qwen3.5-9B: —

Writing & Preference benchmarks
BenchmarkMuse SparkQwen3.5-9B
LMArena Text1474—
LMArena Creative Writing1459—
LMArena Multi-Turn1477—

Frequently asked questions

Is Muse Spark better than Qwen3.5-9B?

Muse Spark is the stronger model overall, scoring 50.6 to 33.8 on the Noometry Index.

Is Muse Spark or Qwen3.5-9B better for coding?

Muse Spark scores higher on coding benchmarks: 46.2 versus 35.9 in the Noometry coding category.

How many benchmarks do Muse Spark and Qwen3.5-9B share?

5 benchmarks have published results for both models. Muse Spark has 27 scored results on Noometry and Qwen3.5-9B has 10.

Related comparisons

Go deeper