Model comparison

Grok 4.6 vs Muse Spark

Grok 4.6 is the stronger model overall, scoring 56.9 to 50.6 on the Noometry Index.

Last verified . 24 shared benchmarks.

Grok 4.6 xAI

56.9

Rank #21 Confirmed

Muse Spark Meta

50.6

Rank #46 Confirmed

Summary

  • They share 24 benchmarks with published results for both. Grok 4.6 scores higher in 5 categories and Muse Spark in 4 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.6 leads 61.4 to 35.9.
  • The biggest single-benchmark swing is ProofBench: 51% for Grok 4.6 and 17% for Muse Spark.

Side by side

Grok 4.6 and Muse Spark specifications
Grok 4.6Muse Spark
ProviderxAIMeta
Noometry Index56.950.6
Released2026-08-122026-04-08
WeightsProprietaryProprietary
Context window500K—
Max output500K—
Input $ / M tokens$2—
Output $ / M tokens$6—
Results tracked4927

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.6 leads

Grok 4.6: 58.5 (#16), Muse Spark: 46.2 (#69)

Coding benchmarks
BenchmarkGrok 4.6Muse Spark
SciCode56.5%51.5%
LMArena Coding14651481
DeepSWE67.5%—
FrontierCode48%—
CursorBench41.4%—
LMArena WebDev1617—
FrontierSWE25.3%—
WeirdML67.3%—
ALE-Bench1,508—

Agentic & Tool Use Not comparable

Grok 4.6: 39.4 (#27), Muse Spark: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.6Muse Spark
APEX-Agents65.3%—
GDP.pdf17.2%—
Vending-Bench 29,047—

Reasoning Grok 4.6 leads

Grok 4.6: 61.4 (#20), Muse Spark: 35.9 (#67)

Reasoning benchmarks
BenchmarkGrok 4.6Muse Spark
CritPt19.7%11.3%
LMArena Hard Prompts14471474
Epoch Capabilities Index156.44152.04
ARC-AGI-267.1%—
SimpleBench75.9%—
NYT Connections (extended)80%—
ARC-AGI-187.5%—
Chess Puzzles40%—
EBR-Bench30.5%—
Mystery Game Puzzles34%—
DTBench97.3%—
LMCA48.5%—

Math Grok 4.6 leads

Grok 4.6: 67.0 (#24), Muse Spark: 47.8 (#66)

Math benchmarks
BenchmarkGrok 4.6Muse Spark
OTIS Mock AIME 2024-202599.2%88.9%
ProofBench51%17%
LMArena Math14231455
FrontierMath (Tiers 1-3)66%—
FrontierMath Tier 431.7%—
FrontierMath (Feb 2025 set)—39%
FrontierMath Tier 4 (v1)—14.6%

Knowledge Muse Spark leads

Grok 4.6: 63.3 (#20), Muse Spark: 65.7 (#13)

Knowledge benchmarks
BenchmarkGrok 4.6Muse Spark
GPQA Diamond94%89.8%
LMArena Expert14671457
Humanity's Last Exam—40.6%
SimpleQA Verified49.3%—

Multimodal Too close to call

Grok 4.6: 43.6 (#23), Muse Spark: 43.4 (#24)

Multimodal benchmarks
BenchmarkGrok 4.6Muse Spark
LMArena Vision12631306
LMArena Document14521444
Blueprint-Bench 233.2%—
Furniture Assembly40%—

Multilingual Muse Spark leads

Grok 4.6: 53.0 (#74), Muse Spark: 56.1 (#24)

Multilingual benchmarks
BenchmarkGrok 4.6Muse Spark
LMArena Non-English14201464
LMArena Chinese14801509
LMArena French14611497
LMArena German14311497
LMArena Korean13971459
LMArena Russian14221466
LMArena Spanish14041472
LMArena Japanese1376—

Instruction Following Too close to call

Grok 4.6: 75.4 (#63), Muse Spark: 75.9 (#51)

Instruction Following benchmarks
BenchmarkGrok 4.6Muse Spark
LMArena Instruction Following14311442

Long Context Too close to call

Grok 4.6: 44.5 (#66), Muse Spark: 44.4 (#69)

Long Context benchmarks
BenchmarkGrok 4.6Muse Spark
LMArena Longer Query14541451

Writing & Preference Muse Spark leads

Grok 4.6: 62.3 (#80), Muse Spark: 66.0 (#39)

Writing & Preference benchmarks
BenchmarkGrok 4.6Muse Spark
LMArena Text14281474
LMArena Creative Writing14281459
LMArena Multi-Turn14251477

Frequently asked questions

Is Grok 4.6 better than Muse Spark?

Grok 4.6 is the stronger model overall, scoring 56.9 to 50.6 on the Noometry Index.

Is Grok 4.6 or Muse Spark better for coding?

Grok 4.6 scores higher on coding benchmarks: 58.5 versus 46.2 in the Noometry coding category.

How many benchmarks do Grok 4.6 and Muse Spark share?

24 benchmarks have published results for both models. Grok 4.6 has 49 scored results on Noometry and Muse Spark has 27.

Related comparisons

Go deeper