Model comparison

Grok 4.3 vs Muse Spark

Muse Spark is the stronger model overall, scoring 50.6 to 43.8 on the Noometry Index.

Last verified . 23 shared benchmarks.

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Muse Spark Meta

50.6

Rank #46 Confirmed

Summary

  • They share 23 benchmarks with published results for both. Grok 4.3 scores higher in 0 categories and Muse Spark in 9 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Muse Spark leads 65.7 to 52.5.
  • The biggest single-benchmark swing is ProofBench: 11% for Grok 4.3 and 17% for Muse Spark.

Side by side

Grok 4.3 and Muse Spark specifications
Grok 4.3Muse Spark
ProviderxAIMeta
Noometry Index43.850.6
Released2026-04-172026-04-08
WeightsProprietaryProprietary
Context window1M—
Max output30K—
Input $ / M tokens$1.25—
Output $ / M tokens$2.50—
Results tracked4027

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Muse Spark leads

Grok 4.3: 41.6 (#121), Muse Spark: 46.2 (#69)

Coding benchmarks
BenchmarkGrok 4.3Muse Spark
SciCode47.3%51.5%
LMArena Coding14151481
LMArena WebDev1357—
WeirdML49.9%—
ALE-Bench944.17—

Agentic & Tool Use Not comparable

Grok 4.3: 27.7 (#99), Muse Spark: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.3Muse Spark
GDP.pdf8%—
LMArena Search1165—
Vending-Bench 235.26—

Reasoning Too close to call

Grok 4.3: 35.9 (#68), Muse Spark: 35.9 (#67)

Reasoning benchmarks
BenchmarkGrok 4.3Muse Spark
CritPt8%11.3%
LMArena Hard Prompts13961474
Epoch Capabilities Index149.16152.04
NYT Connections (extended)55.2%—
Chess Puzzles25%—
DTBench90.7%—
LMCA38.3%—
ForecastBench60.3—

Math Muse Spark leads

Grok 4.3: 46.0 (#74), Muse Spark: 47.8 (#66)

Math benchmarks
BenchmarkGrok 4.3Muse Spark
OTIS Mock AIME 2024-202593.3%88.9%
ProofBench11%17%
LMArena Math13881455
FrontierMath (Tiers 1-3)42.8%—
FrontierMath Tier 414.6%—
FrontierMath (Feb 2025 set)—39%
FrontierMath Tier 4 (v1)—14.6%

Knowledge Muse Spark leads

Grok 4.3: 52.5 (#62), Muse Spark: 65.7 (#13)

Knowledge benchmarks
BenchmarkGrok 4.3Muse Spark
GPQA Diamond88.8%89.8%
LMArena Expert13851457
Humanity's Last Exam—40.6%
SimpleQA Verified33.2%—

Multimodal Muse Spark leads

Grok 4.3: 31.6 (#104), Muse Spark: 43.4 (#24)

Multimodal benchmarks
BenchmarkGrok 4.3Muse Spark
LMArena Vision12291306
Blueprint-Bench 20%—
LMArena Document—1444

Multilingual Muse Spark leads

Grok 4.3: 50.5 (#120), Muse Spark: 56.1 (#24)

Multilingual benchmarks
BenchmarkGrok 4.3Muse Spark
LMArena Non-English13851464
LMArena Chinese14221509
LMArena French14121497
LMArena German13951497
LMArena Korean13561459
LMArena Russian13991466
LMArena Spanish13981472
LMArena Japanese1379—

Instruction Following Muse Spark leads

Grok 4.3: 72.1 (#140), Muse Spark: 75.9 (#51)

Instruction Following benchmarks
BenchmarkGrok 4.3Muse Spark
LMArena Instruction Following13661442

Long Context Muse Spark leads

Grok 4.3: 42.5 (#123), Muse Spark: 44.4 (#69)

Long Context benchmarks
BenchmarkGrok 4.3Muse Spark
LMArena Longer Query13931451

Writing & Preference Muse Spark leads

Grok 4.3: 58.5 (#118), Muse Spark: 66.0 (#39)

Writing & Preference benchmarks
BenchmarkGrok 4.3Muse Spark
LMArena Text13971474
LMArena Creative Writing13801459
LMArena Multi-Turn14061477
EQ-Bench 41075—

Frequently asked questions

Is Grok 4.3 better than Muse Spark?

Muse Spark is the stronger model overall, scoring 50.6 to 43.8 on the Noometry Index.

Is Grok 4.3 or Muse Spark better for coding?

Muse Spark scores higher on coding benchmarks: 46.2 versus 41.6 in the Noometry coding category.

How many benchmarks do Grok 4.3 and Muse Spark share?

23 benchmarks have published results for both models. Grok 4.3 has 40 scored results on Noometry and Muse Spark has 27.

Related comparisons

Go deeper