Model comparison

Muse Spark vs Muse Spark 1.1

Muse Spark and Muse Spark 1.1 score almost the same on the Noometry Index (50.6 vs 49.9), so choose on price, context window or the category you care about most.

Last verified . 22 shared benchmarks.

Muse Spark Meta

50.6

Rank #46 Confirmed

Muse Spark 1.1 Meta

49.9

Rank #51 Confirmed

Summary

  • They share 22 benchmarks with published results for both. Muse Spark scores higher in 3 categories and Muse Spark 1.1 in 6 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Muse Spark leads 65.7 to 53.1.
  • The biggest single-benchmark swing is ProofBench: 17% for Muse Spark and 39% for Muse Spark 1.1.

Side by side

Muse Spark and Muse Spark 1.1 specifications
Muse SparkMuse Spark 1.1
ProviderMetaMeta
Noometry Index50.649.9
Released2026-04-082026-04-08
WeightsProprietaryProprietary
Context window—1.05M
Max output—131K
Input $ / M tokens—$1.25
Output $ / M tokens—$4.25
Results tracked2737

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Muse Spark 1.1 leads

Muse Spark: 46.2 (#69), Muse Spark 1.1: 51.3 (#40)

Coding benchmarks
BenchmarkMuse SparkMuse Spark 1.1
SciCode51.5%58.8%
LMArena Coding14811498
DeepSWE—53.3%
LMArena WebDev—1542

Agentic & Tool Use Not comparable

Muse Spark: —, Muse Spark 1.1: 30.8 (#73)

Agentic & Tool Use benchmarks
BenchmarkMuse SparkMuse Spark 1.1
APEX-Agents—31.8%
τ²-bench Banking—40.5%
GBAEval—7.9%
GDP.pdf—15%
Vending-Bench 2—6,520

Reasoning Muse Spark 1.1 leads

Muse Spark: 35.9 (#67), Muse Spark 1.1: 47.1 (#44)

Reasoning benchmarks
BenchmarkMuse SparkMuse Spark 1.1
CritPt11.3%15.1%
LMArena Hard Prompts14741486
Epoch Capabilities Index152.04154.21
NYT Connections (extended)—84.9%
DTBench—94.4%
LMCA—49.9%
Surface Evolver Bench—52.5%

Math Muse Spark leads

Muse Spark: 47.8 (#66), Muse Spark 1.1: 45.5 (#76)

Math benchmarks
BenchmarkMuse SparkMuse Spark 1.1
ProofBench17%39%
LMArena Math14551483
OTIS Mock AIME 2024-202588.9%—
FrontierMath (Feb 2025 set)39%—
FrontierMath Tier 4 (v1)14.6%—

Knowledge Muse Spark leads

Muse Spark: 65.7 (#13), Muse Spark 1.1: 53.1 (#59)

Knowledge benchmarks
BenchmarkMuse SparkMuse Spark 1.1
LMArena Expert14571478
GPQA Diamond89.8%—
Humanity's Last Exam40.6%—
SimpleQA Verified—57.8%

Multimodal Too close to call

Muse Spark: 43.4 (#24), Muse Spark 1.1: 42.6 (#29)

Multimodal benchmarks
BenchmarkMuse SparkMuse Spark 1.1
LMArena Vision13061293
LMArena Document14441465

Multilingual Too close to call

Muse Spark: 56.1 (#24), Muse Spark 1.1: 56.7 (#17)

Multilingual benchmarks
BenchmarkMuse SparkMuse Spark 1.1
LMArena Non-English14641472
LMArena Chinese15091518
LMArena French14971494
LMArena German14971466
LMArena Korean14591458
LMArena Russian14661483
LMArena Spanish14721464
LMArena Japanese—1451

Instruction Following Too close to call

Muse Spark: 75.9 (#51), Muse Spark 1.1: 76.5 (#39)

Instruction Following benchmarks
BenchmarkMuse SparkMuse Spark 1.1
LMArena Instruction Following14421457

Long Context Too close to call

Muse Spark: 44.4 (#69), Muse Spark 1.1: 44.8 (#58)

Long Context benchmarks
BenchmarkMuse SparkMuse Spark 1.1
LMArena Longer Query14511462

Writing & Preference Muse Spark 1.1 leads

Muse Spark: 66.0 (#39), Muse Spark 1.1: 73.4 (#11)

Writing & Preference benchmarks
BenchmarkMuse SparkMuse Spark 1.1
LMArena Text14741479
LMArena Creative Writing14591437
LMArena Multi-Turn14771485
EQ-Bench Creative Writing—1927
EQ-Bench 4—1260

Frequently asked questions

Is Muse Spark better than Muse Spark 1.1?

Muse Spark and Muse Spark 1.1 score almost the same on the Noometry Index (50.6 vs 49.9), so choose on price, context window or the category you care about most.

Is Muse Spark or Muse Spark 1.1 better for coding?

Muse Spark 1.1 scores higher on coding benchmarks: 51.3 versus 46.2 in the Noometry coding category.

How many benchmarks do Muse Spark and Muse Spark 1.1 share?

22 benchmarks have published results for both models. Muse Spark has 27 scored results on Noometry and Muse Spark 1.1 has 37.

Related comparisons

Go deeper