Model comparison

Claude 3.5 Sonnet vs Muse Spark 1.1

Muse Spark 1.1 is the stronger model overall, scoring 49.9 to 34.6 on the Noometry Index.

Last verified . 21 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

Muse Spark 1.1 Meta

49.9

Rank #51 Confirmed

Summary

  • They share 21 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 1 category and Muse Spark 1.1 in 9 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where Muse Spark 1.1 leads 45.5 to 19.2.
  • The biggest single-benchmark swing is DTBench: 67.8% for Claude 3.5 Sonnet and 94.4% for Muse Spark 1.1.

Side by side

Claude 3.5 Sonnet and Muse Spark 1.1 specifications
Claude 3.5 SonnetMuse Spark 1.1
ProviderAnthropicMeta
Noometry Index34.649.9
Released2024-06-202026-04-08
WeightsProprietaryProprietary
Context window—1.05M
Max output—131K
Input $ / M tokens—$1.25
Output $ / M tokens—$4.25
Results tracked6037

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Muse Spark 1.1 leads

Claude 3.5 Sonnet: 39.0 (#165), Muse Spark 1.1: 51.3 (#40)

Coding benchmarks
BenchmarkClaude 3.5 SonnetMuse Spark 1.1
LMArena Coding13421498
DeepSWE—53.3%
Aider Polyglot51.6%—
LMArena WebDev—1542
SciCode—58.8%
GSO4.6%—
WeirdML40%—
BigCodeBench Instruct46.8%—
LiveBench Coding67.1%—
BigCodeBench Complete58.6%—
CadEval48%—
HumanEval+81.7%—
MBPP+74.3%—

Agentic & Tool Use Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 32.3 (#67), Muse Spark 1.1: 30.8 (#73)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 SonnetMuse Spark 1.1
APEX-Agents—31.8%
TheAgentCompany24%—
τ²-bench Banking—40.5%
Cybench17.5%—
BALROG32.6%—
GBAEval—7.9%
GDP.pdf—15%
METR Time Horizons45.2%—
Vending-Bench 2—6,520

Reasoning Muse Spark 1.1 leads

Claude 3.5 Sonnet: 23.1 (#183), Muse Spark 1.1: 47.1 (#44)

Reasoning benchmarks
BenchmarkClaude 3.5 SonnetMuse Spark 1.1
LMArena Hard Prompts13051486
DTBench67.8%94.4%
Epoch Capabilities Index133.55154.21
SimpleBench41.4%—
NYT Connections (extended)—84.9%
CritPt—15.1%
EnigmaEval0.9%—
LiveBench Reasoning56.7%—
LiveBench Data Analysis55%—
LMCA—49.9%
Surface Evolver Bench—52.5%
ForecastBench60.7—
LiveBench59%—

Math Muse Spark 1.1 leads

Claude 3.5 Sonnet: 19.2 (#288), Muse Spark 1.1: 45.5 (#76)

Math benchmarks
BenchmarkClaude 3.5 SonnetMuse Spark 1.1
LMArena Math13071483
OTIS Mock AIME 2024-20258.5%—
ProofBench—39%
Omni-MATH27.6%—
LiveBench Math52.3%—
MATH Level 556.9%—
FrontierMath (Feb 2025 set)2.1%—
FrontierMath Tier 4 (v1)0%—

Knowledge Muse Spark 1.1 leads

Claude 3.5 Sonnet: 28.6 (#245), Muse Spark 1.1: 53.1 (#59)

Knowledge benchmarks
BenchmarkClaude 3.5 SonnetMuse Spark 1.1
LMArena Expert12651478
GPQA Diamond55.3%—
Humanity's Last Exam4.1%—
SimpleQA Verified—57.8%
MMLU-Pro77.7%—
Confabulations19.9%—
GPQA (HELM)56.5%—
MMLU87.3%—

Multimodal Muse Spark 1.1 leads

Claude 3.5 Sonnet: 26.5 (#120), Muse Spark 1.1: 42.6 (#29)

Multimodal benchmarks
BenchmarkClaude 3.5 SonnetMuse Spark 1.1
LMArena Vision11251293
Video-MME60%—
GeoBench62%—
VPCT33%—
LMArena Document—1465

Multilingual Muse Spark 1.1 leads

Claude 3.5 Sonnet: 43.2 (#185), Muse Spark 1.1: 56.7 (#17)

Multilingual benchmarks
BenchmarkClaude 3.5 SonnetMuse Spark 1.1
LMArena Non-English12831472
LMArena Chinese12721518
LMArena French13051494
LMArena German12971466
LMArena Japanese12341451
LMArena Korean12001458
LMArena Russian13061483
LMArena Spanish12901464

Instruction Following Muse Spark 1.1 leads

Claude 3.5 Sonnet: 68.8 (#182), Muse Spark 1.1: 76.5 (#39)

Instruction Following benchmarks
BenchmarkClaude 3.5 SonnetMuse Spark 1.1
LMArena Instruction Following12971457
LiveBench Instruction Following69.3%—
IFEval85.5%—

Long Context Muse Spark 1.1 leads

Claude 3.5 Sonnet: 39.9 (#167), Muse Spark 1.1: 44.8 (#58)

Long Context benchmarks
BenchmarkClaude 3.5 SonnetMuse Spark 1.1
LMArena Longer Query13111462

Writing & Preference Muse Spark 1.1 leads

Claude 3.5 Sonnet: 52.9 (#164), Muse Spark 1.1: 73.4 (#11)

Writing & Preference benchmarks
BenchmarkClaude 3.5 SonnetMuse Spark 1.1
LMArena Text12981479
LMArena Creative Writing12921437
EQ-Bench Creative Writing14511927
LMArena Multi-Turn13261485
Short-Story Creative Writing80.3%—
WildBench79.2%—
EQ-Bench 4—1260
LiveBench Language53.8%—

Frequently asked questions

Is Claude 3.5 Sonnet better than Muse Spark 1.1?

Muse Spark 1.1 is the stronger model overall, scoring 49.9 to 34.6 on the Noometry Index.

Is Claude 3.5 Sonnet or Muse Spark 1.1 better for coding?

Muse Spark 1.1 scores higher on coding benchmarks: 51.3 versus 39.0 in the Noometry coding category.

How many benchmarks do Claude 3.5 Sonnet and Muse Spark 1.1 share?

21 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and Muse Spark 1.1 has 37.

Related comparisons

Go deeper