Model comparison

Muse Spark 1.2 vs Qwen Max

Muse Spark 1.2 is the stronger model overall, scoring 50.3 to 34.7 on the Noometry Index.

Last verified . 14 shared benchmarks.

Muse Spark 1.2 Meta

50.3

Rank #48 Confirmed

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 14 benchmarks with published results for both. Muse Spark 1.2 scores higher in 8 categories and Qwen Max in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Muse Spark 1.2 leads 51.3 to 25.1.
  • Muse Spark 1.2 is cheaper at $1.25 / $4.25 per million input/output tokens, against $1.60 / $6.40 for Qwen Max.
  • Muse Spark 1.2 accepts more context: 1.05M tokens versus 33K.

Side by side

Muse Spark 1.2 and Qwen Max specifications
Muse Spark 1.2Qwen Max
ProviderMetaAlibaba (Qwen)
Noometry Index50.334.7
Released2026-08-052024-04-03
WeightsProprietaryProprietary
Context window1.05M33K
Max output131K8K
Input $ / M tokens$1.25$1.60
Output $ / M tokens$4.25$6.40
Results tracked3123

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Muse Spark 1.2 leads

Muse Spark 1.2: 49.2 (#51), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkMuse Spark 1.2Qwen Max
LMArena Coding14951288
DeepSWE54.9%—
Aider Polyglot—21.8%
LMArena WebDev1533—
FrontierSWE12%—
SciCode56.4%—
WeirdML60.3%—

Agentic & Tool Use Not comparable

Muse Spark 1.2: 29.4 (#87), Qwen Max: —

Agentic & Tool Use benchmarks
BenchmarkMuse Spark 1.2Qwen Max
APEX-Agents36.4%—
GDP.pdf16%—

Reasoning Muse Spark 1.2 leads

Muse Spark 1.2: 51.3 (#34), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkMuse Spark 1.2Qwen Max
LMArena Hard Prompts14861269
SimpleBench74.5%—
NYT Connections (extended)79.2%—
CritPt17.7%—
DTBench94.7%—
LMCA48.4%—
Epoch Capabilities Index154.87—

Math Muse Spark 1.2 leads

Muse Spark 1.2: 46.4 (#70), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkMuse Spark 1.2Qwen Max
LMArena Math14711275
OTIS Mock AIME 2024-2025—16.1%
ProofBench43%—
MATH Level 5—67.2%
FrontierMath (Feb 2025 set)—1%

Knowledge Muse Spark 1.2 leads

Muse Spark 1.2: 54.1 (#53), Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkMuse Spark 1.2Qwen Max
LMArena Expert14801248
GPQA Diamond—56.1%
SimpleQA Verified60.3%—

Multimodal Not comparable

Muse Spark 1.2: 43.4 (#25), Qwen Max: —

Multimodal benchmarks
BenchmarkMuse Spark 1.2Qwen Max
LMArena Vision1305—

Multilingual Muse Spark 1.2 leads

Muse Spark 1.2: 57.1 (#11), Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkMuse Spark 1.2Qwen Max
LMArena Non-English14781263
LMArena Chinese15111254
LMArena French15131330
LMArena Russian14871274
LMArena Spanish14981290
LMArena German—1254
LMArena Japanese—1205
LMArena Korean—1142

Instruction Following Muse Spark 1.2 leads

Muse Spark 1.2: 76.7 (#36), Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkMuse Spark 1.2Qwen Max
LMArena Instruction Following14611262

Long Context Muse Spark 1.2 leads

Muse Spark 1.2: 45.2 (#48), Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkMuse Spark 1.2Qwen Max
LMArena Longer Query14751288
Fiction.LiveBench—66.7%

Writing & Preference Muse Spark 1.2 leads

Muse Spark 1.2: 72.3 (#14), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkMuse Spark 1.2Qwen Max
LMArena Text14821282
LMArena Creative Writing14491248
LMArena Multi-Turn14941277
EQ-Bench Creative Writing1840—

Frequently asked questions

Is Muse Spark 1.2 better than Qwen Max?

Muse Spark 1.2 is the stronger model overall, scoring 50.3 to 34.7 on the Noometry Index.

Which is cheaper, Muse Spark 1.2 or Qwen Max?

Muse Spark 1.2 is cheaper. It lists at $1.25 per million input tokens and $4.25 per million output tokens; Qwen Max lists at $1.60 and $6.40.

Is Muse Spark 1.2 or Qwen Max better for coding?

Muse Spark 1.2 scores higher on coding benchmarks: 49.2 versus 30.7 in the Noometry coding category.

Which has the bigger context window?

Muse Spark 1.2 does, with 1.05M tokens against 33K.

How many benchmarks do Muse Spark 1.2 and Qwen Max share?

14 benchmarks have published results for both models. Muse Spark 1.2 has 31 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper