Model comparison

Grok 4.3 vs Muse Spark 1.2

Muse Spark 1.2 is the stronger model overall, scoring 50.3 to 43.8 on the Noometry Index.

Last verified . 26 shared benchmarks.

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Muse Spark 1.2 Meta

50.3

Rank #48 Confirmed

Summary

  • They share 26 benchmarks with published results for both. Grok 4.3 scores higher in 0 categories and Muse Spark 1.2 in 10 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Muse Spark 1.2 leads 51.3 to 35.9.
  • The biggest single-benchmark swing is ProofBench: 11% for Grok 4.3 and 43% for Muse Spark 1.2.
  • Grok 4.3 is cheaper at $1.25 / $2.50 per million input/output tokens, against $1.25 / $4.25 for Muse Spark 1.2.
  • Muse Spark 1.2 accepts more context: 1.05M tokens versus 1M.

Side by side

Grok 4.3 and Muse Spark 1.2 specifications
Grok 4.3Muse Spark 1.2
ProviderxAIMeta
Noometry Index43.850.3
Released2026-04-172026-08-05
WeightsProprietaryProprietary
Context window1M1.05M
Max output30K131K
Input $ / M tokens$1.25$1.25
Output $ / M tokens$2.50$4.25
Results tracked4031

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Muse Spark 1.2 leads

Grok 4.3: 41.6 (#121), Muse Spark 1.2: 49.2 (#51)

Coding benchmarks
BenchmarkGrok 4.3Muse Spark 1.2
LMArena WebDev13571533
SciCode47.3%56.4%
WeirdML49.9%60.3%
LMArena Coding14151495
DeepSWE—54.9%
FrontierSWE—12%
ALE-Bench944.17—

Agentic & Tool Use Muse Spark 1.2 leads

Grok 4.3: 27.7 (#99), Muse Spark 1.2: 29.4 (#87)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.3Muse Spark 1.2
GDP.pdf8%16%
APEX-Agents—36.4%
LMArena Search1165—
Vending-Bench 235.26—

Reasoning Muse Spark 1.2 leads

Grok 4.3: 35.9 (#68), Muse Spark 1.2: 51.3 (#34)

Reasoning benchmarks
BenchmarkGrok 4.3Muse Spark 1.2
NYT Connections (extended)55.2%79.2%
CritPt8%17.7%
LMArena Hard Prompts13961486
DTBench90.7%94.7%
LMCA38.3%48.4%
Epoch Capabilities Index149.16154.87
SimpleBench—74.5%
Chess Puzzles25%—
ForecastBench60.3—

Math Too close to call

Grok 4.3: 46.0 (#74), Muse Spark 1.2: 46.4 (#70)

Math benchmarks
BenchmarkGrok 4.3Muse Spark 1.2
ProofBench11%43%
LMArena Math13881471
FrontierMath (Tiers 1-3)42.8%—
FrontierMath Tier 414.6%—
OTIS Mock AIME 2024-202593.3%—

Knowledge Muse Spark 1.2 leads

Grok 4.3: 52.5 (#62), Muse Spark 1.2: 54.1 (#53)

Knowledge benchmarks
BenchmarkGrok 4.3Muse Spark 1.2
SimpleQA Verified33.2%60.3%
LMArena Expert13851480
GPQA Diamond88.8%—

Multimodal Muse Spark 1.2 leads

Grok 4.3: 31.6 (#104), Muse Spark 1.2: 43.4 (#25)

Multimodal benchmarks
BenchmarkGrok 4.3Muse Spark 1.2
LMArena Vision12291305
Blueprint-Bench 20%—

Multilingual Muse Spark 1.2 leads

Grok 4.3: 50.5 (#120), Muse Spark 1.2: 57.1 (#11)

Multilingual benchmarks
BenchmarkGrok 4.3Muse Spark 1.2
LMArena Non-English13851478
LMArena Chinese14221511
LMArena French14121513
LMArena Russian13991487
LMArena Spanish13981498
LMArena German1395—
LMArena Japanese1379—
LMArena Korean1356—

Instruction Following Muse Spark 1.2 leads

Grok 4.3: 72.1 (#140), Muse Spark 1.2: 76.7 (#36)

Instruction Following benchmarks
BenchmarkGrok 4.3Muse Spark 1.2
LMArena Instruction Following13661461

Long Context Muse Spark 1.2 leads

Grok 4.3: 42.5 (#123), Muse Spark 1.2: 45.2 (#48)

Long Context benchmarks
BenchmarkGrok 4.3Muse Spark 1.2
LMArena Longer Query13931475

Writing & Preference Muse Spark 1.2 leads

Grok 4.3: 58.5 (#118), Muse Spark 1.2: 72.3 (#14)

Writing & Preference benchmarks
BenchmarkGrok 4.3Muse Spark 1.2
LMArena Text13971482
LMArena Creative Writing13801449
LMArena Multi-Turn14061494
EQ-Bench Creative Writing—1840
EQ-Bench 41075—

Frequently asked questions

Is Grok 4.3 better than Muse Spark 1.2?

Muse Spark 1.2 is the stronger model overall, scoring 50.3 to 43.8 on the Noometry Index.

Which is cheaper, Grok 4.3 or Muse Spark 1.2?

Grok 4.3 is cheaper. It lists at $1.25 per million input tokens and $2.50 per million output tokens; Muse Spark 1.2 lists at $1.25 and $4.25.

Is Grok 4.3 or Muse Spark 1.2 better for coding?

Muse Spark 1.2 scores higher on coding benchmarks: 49.2 versus 41.6 in the Noometry coding category.

Which has the bigger context window?

Muse Spark 1.2 does, with 1.05M tokens against 1M.

How many benchmarks do Grok 4.3 and Muse Spark 1.2 share?

26 benchmarks have published results for both models. Grok 4.3 has 40 scored results on Noometry and Muse Spark 1.2 has 31.

Related comparisons

Go deeper