Model comparison

Grok 4.1 vs Muse Spark

Muse Spark is the stronger model overall, scoring 50.6 to 41.5 on the Noometry Index.

Last verified . 16 shared benchmarks.

Grok 4.1 xAI

41.5

Rank #134 Confirmed

Muse Spark Meta

50.6

Rank #46 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Grok 4.1 scores higher in 0 categories and Muse Spark in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Muse Spark leads 65.7 to 39.5.

Side by side

Grok 4.1 and Muse Spark specifications
Grok 4.1Muse Spark
ProviderxAIMeta
Noometry Index41.550.6
Released2025-11-172026-04-08
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1927

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Muse Spark leads

Grok 4.1: 33.7 (#253), Muse Spark: 46.2 (#69)

Coding benchmarks
BenchmarkGrok 4.1Muse Spark
LMArena Coding14451481
LMArena WebDev1214—
SciCode—51.5%

Agentic & Tool Use Not comparable

Grok 4.1: 34.1 (#49), Muse Spark: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.1Muse Spark
Cybench39%—

Reasoning Muse Spark leads

Grok 4.1: 29.5 (#91), Muse Spark: 35.9 (#67)

Reasoning benchmarks
BenchmarkGrok 4.1Muse Spark
LMArena Hard Prompts14351474
CritPt—11.3%
Epoch Capabilities Index—152.04

Math Muse Spark leads

Grok 4.1: 38.9 (#120), Muse Spark: 47.8 (#66)

Math benchmarks
BenchmarkGrok 4.1Muse Spark
LMArena Math14221455
OTIS Mock AIME 2024-2025—88.9%
ProofBench—17%
FrontierMath (Feb 2025 set)—39%
FrontierMath Tier 4 (v1)—14.6%

Knowledge Muse Spark leads

Grok 4.1: 39.5 (#133), Muse Spark: 65.7 (#13)

Knowledge benchmarks
BenchmarkGrok 4.1Muse Spark
LMArena Expert14171457
GPQA Diamond—89.8%
Humanity's Last Exam—40.6%

Multimodal Not comparable

Grok 4.1: —, Muse Spark: 43.4 (#24)

Multimodal benchmarks
BenchmarkGrok 4.1Muse Spark
LMArena Vision—1306
LMArena Document—1444

Multilingual Muse Spark leads

Grok 4.1: 53.4 (#68), Muse Spark: 56.1 (#24)

Multilingual benchmarks
BenchmarkGrok 4.1Muse Spark
LMArena Non-English14251464
LMArena Chinese14651509
LMArena French14481497
LMArena German14461497
LMArena Korean14071459
LMArena Russian14341466
LMArena Spanish14381472
LMArena Japanese1397—

Instruction Following Muse Spark leads

Grok 4.1: 73.8 (#111), Muse Spark: 75.9 (#51)

Instruction Following benchmarks
BenchmarkGrok 4.1Muse Spark
LMArena Instruction Following14001442

Long Context Muse Spark leads

Grok 4.1: 43.2 (#100), Muse Spark: 44.4 (#69)

Long Context benchmarks
BenchmarkGrok 4.1Muse Spark
LMArena Longer Query14161451

Writing & Preference Muse Spark leads

Grok 4.1: 62.4 (#75), Muse Spark: 66.0 (#39)

Writing & Preference benchmarks
BenchmarkGrok 4.1Muse Spark
LMArena Text14371474
LMArena Creative Writing14111459
LMArena Multi-Turn14371477

Frequently asked questions

Is Grok 4.1 better than Muse Spark?

Muse Spark is the stronger model overall, scoring 50.6 to 41.5 on the Noometry Index.

Is Grok 4.1 or Muse Spark better for coding?

Muse Spark scores higher on coding benchmarks: 46.2 versus 33.7 in the Noometry coding category.

How many benchmarks do Grok 4.1 and Muse Spark share?

16 benchmarks have published results for both models. Grok 4.1 has 19 scored results on Noometry and Muse Spark has 27.

Related comparisons

Go deeper