Model comparison

Grok 4.6 vs Muse Spark 1.1

Grok 4.6 is the stronger model overall, scoring 56.9 to 49.9 on the Noometry Index. Muse Spark 1.1 costs 1.5× less per token, which makes it the better buy when Grok 4.6's lead doesn't matter for your workload.

Last verified . 32 shared benchmarks.

Grok 4.6 xAI

56.9

Rank #21 Confirmed

Muse Spark 1.1 Meta

49.9

Rank #51 Confirmed

Summary

  • They share 32 benchmarks with published results for both. Grok 4.6 scores higher in 6 categories and Muse Spark 1.1 in 4 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Grok 4.6 leads 67.0 to 45.5.
  • The biggest single-benchmark swing is APEX-Agents: 65.3% for Grok 4.6 and 31.8% for Muse Spark 1.1.
  • Muse Spark 1.1 is cheaper at $1.25 / $4.25 per million input/output tokens, against $2 / $6 for Grok 4.6.
  • Muse Spark 1.1 accepts more context: 1.05M tokens versus 500K.

Side by side

Grok 4.6 and Muse Spark 1.1 specifications
Grok 4.6Muse Spark 1.1
ProviderxAIMeta
Noometry Index56.949.9
Released2026-08-122026-04-08
WeightsProprietaryProprietary
Context window500K1.05M
Max output500K131K
Input $ / M tokens$2$1.25
Output $ / M tokens$6$4.25
Results tracked4937

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.6 leads

Grok 4.6: 58.5 (#16), Muse Spark 1.1: 51.3 (#40)

Coding benchmarks
BenchmarkGrok 4.6Muse Spark 1.1
DeepSWE67.5%53.3%
LMArena WebDev16171542
SciCode56.5%58.8%
LMArena Coding14651498
FrontierCode48%—
CursorBench41.4%—
FrontierSWE25.3%—
WeirdML67.3%—
ALE-Bench1,508—

Agentic & Tool Use Grok 4.6 leads

Grok 4.6: 39.4 (#27), Muse Spark 1.1: 30.8 (#73)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.6Muse Spark 1.1
APEX-Agents65.3%31.8%
GDP.pdf17.2%15%
Vending-Bench 29,0476,520
τ²-bench Banking—40.5%
GBAEval—7.9%

Reasoning Grok 4.6 leads

Grok 4.6: 61.4 (#20), Muse Spark 1.1: 47.1 (#44)

Reasoning benchmarks
BenchmarkGrok 4.6Muse Spark 1.1
NYT Connections (extended)80%84.9%
CritPt19.7%15.1%
LMArena Hard Prompts14471486
DTBench97.3%94.4%
LMCA48.5%49.9%
Epoch Capabilities Index156.44154.21
ARC-AGI-267.1%—
SimpleBench75.9%—
ARC-AGI-187.5%—
Chess Puzzles40%—
EBR-Bench30.5%—
Mystery Game Puzzles34%—
Surface Evolver Bench—52.5%

Math Grok 4.6 leads

Grok 4.6: 67.0 (#24), Muse Spark 1.1: 45.5 (#76)

Math benchmarks
BenchmarkGrok 4.6Muse Spark 1.1
ProofBench51%39%
LMArena Math14231483
FrontierMath (Tiers 1-3)66%—
FrontierMath Tier 431.7%—
OTIS Mock AIME 2024-202599.2%—

Knowledge Grok 4.6 leads

Grok 4.6: 63.3 (#20), Muse Spark 1.1: 53.1 (#59)

Knowledge benchmarks
BenchmarkGrok 4.6Muse Spark 1.1
SimpleQA Verified49.3%57.8%
LMArena Expert14671478
GPQA Diamond94%—

Multimodal Too close to call

Grok 4.6: 43.6 (#23), Muse Spark 1.1: 42.6 (#29)

Multimodal benchmarks
BenchmarkGrok 4.6Muse Spark 1.1
LMArena Vision12631293
LMArena Document14521465
Blueprint-Bench 233.2%—
Furniture Assembly40%—

Multilingual Muse Spark 1.1 leads

Grok 4.6: 53.0 (#74), Muse Spark 1.1: 56.7 (#17)

Multilingual benchmarks
BenchmarkGrok 4.6Muse Spark 1.1
LMArena Non-English14201472
LMArena Chinese14801518
LMArena French14611494
LMArena German14311466
LMArena Japanese13761451
LMArena Korean13971458
LMArena Russian14221483
LMArena Spanish14041464

Instruction Following Muse Spark 1.1 leads

Grok 4.6: 75.4 (#63), Muse Spark 1.1: 76.5 (#39)

Instruction Following benchmarks
BenchmarkGrok 4.6Muse Spark 1.1
LMArena Instruction Following14311457

Long Context Too close to call

Grok 4.6: 44.5 (#66), Muse Spark 1.1: 44.8 (#58)

Long Context benchmarks
BenchmarkGrok 4.6Muse Spark 1.1
LMArena Longer Query14541462

Writing & Preference Muse Spark 1.1 leads

Grok 4.6: 62.3 (#80), Muse Spark 1.1: 73.4 (#11)

Writing & Preference benchmarks
BenchmarkGrok 4.6Muse Spark 1.1
LMArena Text14281479
LMArena Creative Writing14281437
LMArena Multi-Turn14251485
EQ-Bench Creative Writing—1927
EQ-Bench 4—1260

Frequently asked questions

Is Grok 4.6 better than Muse Spark 1.1?

Grok 4.6 is the stronger model overall, scoring 56.9 to 49.9 on the Noometry Index. Muse Spark 1.1 costs 1.5× less per token, which makes it the better buy when Grok 4.6's lead doesn't matter for your workload.

Which is cheaper, Grok 4.6 or Muse Spark 1.1?

Muse Spark 1.1 is cheaper. It lists at $1.25 per million input tokens and $4.25 per million output tokens; Grok 4.6 lists at $2 and $6.

Is Grok 4.6 or Muse Spark 1.1 better for coding?

Grok 4.6 scores higher on coding benchmarks: 58.5 versus 51.3 in the Noometry coding category.

Which has the bigger context window?

Muse Spark 1.1 does, with 1.05M tokens against 500K.

How many benchmarks do Grok 4.6 and Muse Spark 1.1 share?

32 benchmarks have published results for both models. Grok 4.6 has 49 scored results on Noometry and Muse Spark 1.1 has 37.

Related comparisons

Go deeper