Model comparison

Muse Spark 1.1 vs Qwen Turbo

Muse Spark 1.1 is the stronger model overall, scoring 49.9 to 27.1 on the Noometry Index. Qwen Turbo costs 23× less per token, which makes it the better buy when Muse Spark 1.1's lead doesn't matter for your workload.

Last verified . 0 shared benchmarks.

Muse Spark 1.1 Meta

49.9

Rank #51 Confirmed

Qwen Turbo Alibaba (Qwen)

27.1

Rank #335 Reported

Summary

  • The widest gap is in knowledge, where Muse Spark 1.1 leads 53.1 to 22.2.
  • Qwen Turbo is cheaper at $0.05 / $0.20 per million input/output tokens, against $1.25 / $4.25 for Muse Spark 1.1.
  • Muse Spark 1.1 accepts more context: 1.05M tokens versus 1M.

Side by side

Muse Spark 1.1 and Qwen Turbo specifications
Muse Spark 1.1Qwen Turbo
ProviderMetaAlibaba (Qwen)
Noometry Index49.927.1
Released2026-04-082024-11-01
WeightsProprietaryProprietary
Context window1.05M1M
Max output131K16K
Input $ / M tokens$1.25$0.05
Output $ / M tokens$4.25$0.20
Results tracked373

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Muse Spark 1.1: 51.3 (#40), Qwen Turbo: —

Coding benchmarks
BenchmarkMuse Spark 1.1Qwen Turbo
DeepSWE53.3%—
LMArena WebDev1542—
SciCode58.8%—
LMArena Coding1498—

Agentic & Tool Use Not comparable

Muse Spark 1.1: 30.8 (#73), Qwen Turbo: —

Agentic & Tool Use benchmarks
BenchmarkMuse Spark 1.1Qwen Turbo
APEX-Agents31.8%—
τ²-bench Banking40.5%—
GBAEval7.9%—
GDP.pdf15%—
Vending-Bench 26,520—

Reasoning Not comparable

Muse Spark 1.1: 47.1 (#44), Qwen Turbo: —

Reasoning benchmarks
BenchmarkMuse Spark 1.1Qwen Turbo
NYT Connections (extended)84.9%—
CritPt15.1%—
LMArena Hard Prompts1486—
DTBench94.4%—
LMCA49.9%—
Surface Evolver Bench52.5%—
Epoch Capabilities Index154.21—

Math Muse Spark 1.1 leads

Muse Spark 1.1: 45.5 (#76), Qwen Turbo: 15.3 (#297)

Math benchmarks
BenchmarkMuse Spark 1.1Qwen Turbo
OTIS Mock AIME 2024-2025—6.1%
ProofBench39%—
LMArena Math1483—
MATH Level 5—56.2%

Knowledge Muse Spark 1.1 leads

Muse Spark 1.1: 53.1 (#59), Qwen Turbo: 22.2 (#272)

Knowledge benchmarks
BenchmarkMuse Spark 1.1Qwen Turbo
GPQA Diamond—41.8%
SimpleQA Verified57.8%—
LMArena Expert1478—

Multimodal Not comparable

Muse Spark 1.1: 42.6 (#29), Qwen Turbo: —

Multimodal benchmarks
BenchmarkMuse Spark 1.1Qwen Turbo
LMArena Vision1293—
LMArena Document1465—

Multilingual Not comparable

Muse Spark 1.1: 56.7 (#17), Qwen Turbo: —

Multilingual benchmarks
BenchmarkMuse Spark 1.1Qwen Turbo
LMArena Non-English1472—
LMArena Chinese1518—
LMArena French1494—
LMArena German1466—
LMArena Japanese1451—
LMArena Korean1458—
LMArena Russian1483—
LMArena Spanish1464—

Instruction Following Not comparable

Muse Spark 1.1: 76.5 (#39), Qwen Turbo: —

Instruction Following benchmarks
BenchmarkMuse Spark 1.1Qwen Turbo
LMArena Instruction Following1457—

Long Context Not comparable

Muse Spark 1.1: 44.8 (#58), Qwen Turbo: —

Long Context benchmarks
BenchmarkMuse Spark 1.1Qwen Turbo
LMArena Longer Query1462—

Writing & Preference Not comparable

Muse Spark 1.1: 73.4 (#11), Qwen Turbo: —

Writing & Preference benchmarks
BenchmarkMuse Spark 1.1Qwen Turbo
LMArena Text1479—
LMArena Creative Writing1437—
EQ-Bench Creative Writing1927—
EQ-Bench 41260—
LMArena Multi-Turn1485—

Frequently asked questions

Is Muse Spark 1.1 better than Qwen Turbo?

Muse Spark 1.1 is the stronger model overall, scoring 49.9 to 27.1 on the Noometry Index. Qwen Turbo costs 23× less per token, which makes it the better buy when Muse Spark 1.1's lead doesn't matter for your workload.

Which is cheaper, Muse Spark 1.1 or Qwen Turbo?

Qwen Turbo is cheaper. It lists at $0.05 per million input tokens and $0.20 per million output tokens; Muse Spark 1.1 lists at $1.25 and $4.25.

Which has the bigger context window?

Muse Spark 1.1 does, with 1.05M tokens against 1M.

How many benchmarks do Muse Spark 1.1 and Qwen Turbo share?

0 benchmarks have published results for both models. Muse Spark 1.1 has 37 scored results on Noometry and Qwen Turbo has 3.

Related comparisons

Go deeper