Model comparison

Grok 4 vs Muse Spark 1.1

Muse Spark 1.1 is the stronger model overall, scoring 49.9 to 48.1 on the Noometry Index.

Last verified . 19 shared benchmarks.

Grok 4 xAI

48.1

Rank #56 Confirmed

Muse Spark 1.1 Meta

49.9

Rank #51 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Grok 4 scores higher in 5 categories and Muse Spark 1.1 in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Grok 4 leads 63.1 to 44.8.

Side by side

Grok 4 and Muse Spark 1.1 specifications
Grok 4Muse Spark 1.1
ProviderxAIMeta
Noometry Index48.149.9
Released2025-07-092026-04-08
WeightsProprietaryProprietary
Context window—1.05M
Max output—131K
Input $ / M tokens—$1.25
Output $ / M tokens—$4.25
Results tracked4837

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Grok 4: 50.3 (#46), Muse Spark 1.1: 51.3 (#40)

Coding benchmarks
BenchmarkGrok 4Muse Spark 1.1
LMArena Coding14081498
DeepSWE—53.3%
Aider Polyglot79.6%—
LMArena WebDev—1542
SciCode—58.8%
WeirdML45.7%—

Agentic & Tool Use Grok 4 leads

Grok 4: 32.3 (#68), Muse Spark 1.1: 30.8 (#73)

Agentic & Tool Use benchmarks
BenchmarkGrok 4Muse Spark 1.1
Terminal-Bench27.2%—
APEX-Agents—31.8%
Berkeley Function Calling Leaderboard63%—
GDPval21.1%—
τ²-bench Banking—40.5%
Cybench43%—
DeepResearch Bench47.3%—
BALROG43.6%—
GBAEval—7.9%
GDP.pdf—15%
LMArena Search1142—
METR Time Horizons66.6%—
Vending-Bench 2—6,520

Reasoning Muse Spark 1.1 leads

Grok 4: 36.7 (#65), Muse Spark 1.1: 47.1 (#44)

Reasoning benchmarks
BenchmarkGrok 4Muse Spark 1.1
LMArena Hard Prompts14091486
Epoch Capabilities Index146.44154.21
ARC-AGI-216%—
SimpleBench60.5%—
Kagi LLM Benchmark73.6%—
NYT Connections (extended)—84.9%
ARC-AGI-166.7%—
CritPt—15.1%
Chess Puzzles28%—
DTBench—94.4%
LMCA—49.9%
Surface Evolver Bench—52.5%
ForecastBench60.9—

Math Grok 4 leads

Grok 4: 48.4 (#64), Muse Spark 1.1: 45.5 (#76)

Math benchmarks
BenchmarkGrok 4Muse Spark 1.1
LMArena Math14221483
OTIS Mock AIME 2024-202584%—
ProofBench—39%
Omni-MATH60.3%—
FrontierMath (Feb 2025 set)19.7%—
FrontierMath Tier 4 (v1)2.1%—

Knowledge Too close to call

Grok 4: 53.8 (#55), Muse Spark 1.1: 53.1 (#59)

Knowledge benchmarks
BenchmarkGrok 4Muse Spark 1.1
LMArena Expert14151478
GPQA Diamond87%—
SimpleQA Verified—57.8%
MMLU-Pro85.1%—
Confabulations12.4%—
GPQA (HELM)72.7%—

Multimodal Muse Spark 1.1 leads

Grok 4: 33.7 (#94), Muse Spark 1.1: 42.6 (#29)

Multimodal benchmarks
BenchmarkGrok 4Muse Spark 1.1
LMArena Vision12101293
GeoBench45%—
LMArena Document—1465

Multilingual Muse Spark 1.1 leads

Grok 4: 51.8 (#103), Muse Spark 1.1: 56.7 (#17)

Multilingual benchmarks
BenchmarkGrok 4Muse Spark 1.1
LMArena Non-English14031472
LMArena Chinese14271518
LMArena French14181494
LMArena German14291466
LMArena Japanese13941451
LMArena Korean13771458
LMArena Russian14101483
LMArena Spanish14201464

Instruction Following Grok 4 leads

Grok 4: 79.2 (#5), Muse Spark 1.1: 76.5 (#39)

Instruction Following benchmarks
BenchmarkGrok 4Muse Spark 1.1
LMArena Instruction Following13871457
IFEval94.9%—

Long Context Grok 4 leads

Grok 4: 63.1 (#4), Muse Spark 1.1: 44.8 (#58)

Long Context benchmarks
BenchmarkGrok 4Muse Spark 1.1
LMArena Longer Query14091462
Fiction.LiveBench94.4%—

Writing & Preference Muse Spark 1.1 leads

Grok 4: 58.5 (#116), Muse Spark 1.1: 73.4 (#11)

Writing & Preference benchmarks
BenchmarkGrok 4Muse Spark 1.1
LMArena Text14111479
LMArena Creative Writing13971437
LMArena Multi-Turn14161485
Short-Story Creative Writing76.9%—
EQ-Bench Creative Writing—1927
WildBench79.7%—
EQ-Bench 4—1260

Frequently asked questions

Is Grok 4 better than Muse Spark 1.1?

Muse Spark 1.1 is the stronger model overall, scoring 49.9 to 48.1 on the Noometry Index.

Is Grok 4 or Muse Spark 1.1 better for coding?

They score almost the same on coding (50.3 vs 51.3); test both on your own repository before choosing.

How many benchmarks do Grok 4 and Muse Spark 1.1 share?

19 benchmarks have published results for both models. Grok 4 has 48 scored results on Noometry and Muse Spark 1.1 has 37.

Related comparisons

Go deeper