Model comparison

GPT-5.4 mini vs Muse Spark

Muse Spark is the stronger model overall, scoring 50.6 to 45.0 on the Noometry Index.

Last verified . 25 shared benchmarks.

GPT-5.4 mini OpenAI

45.0

Rank #76 Confirmed

Muse Spark Meta

50.6

Rank #46 Confirmed

Summary

  • They share 25 benchmarks with published results for both. GPT-5.4 mini scores higher in 0 categories and Muse Spark in 9 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Muse Spark leads 65.7 to 51.5.

Side by side

GPT-5.4 mini and Muse Spark specifications
GPT-5.4 miniMuse Spark
ProviderOpenAIMeta
Noometry Index45.050.6
Released2026-03-172026-04-08
WeightsProprietaryProprietary
Context window400K—
Max output128K—
Input $ / M tokens$0.75—
Output $ / M tokens$4.50—
Results tracked4627

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

GPT-5.4 mini: 45.2 (#72), Muse Spark: 46.2 (#69)

Coding benchmarks
BenchmarkGPT-5.4 miniMuse Spark
SciCode49.9%51.5%
LMArena Coding14381481
FrontierCode27%—
LMArena WebDev1397—
WeirdML60.3%—
ALE-Bench1,189—

Agentic & Tool Use Not comparable

GPT-5.4 mini: 29.9 (#81), Muse Spark: —

Agentic & Tool Use benchmarks
BenchmarkGPT-5.4 miniMuse Spark
DeepResearch Bench36.3%—

Reasoning Muse Spark leads

GPT-5.4 mini: 30.4 (#85), Muse Spark: 35.9 (#67)

Reasoning benchmarks
BenchmarkGPT-5.4 miniMuse Spark
CritPt10%11.3%
LMArena Hard Prompts14241474
Epoch Capabilities Index148.84152.04
ARC-AGI-218.9%—
Kagi LLM Benchmark37.9%—
NYT Connections (extended)61.8%—
ARC-AGI-163.7%—
Chess Puzzles24%—
Thematic Generalization61.7%—
Mystery Game Puzzles11%—
DTBench80%—
LMCA40.8%—
ForecastBench57—

Math Muse Spark leads

GPT-5.4 mini: 45.5 (#75), Muse Spark: 47.8 (#66)

Math benchmarks
BenchmarkGPT-5.4 miniMuse Spark
OTIS Mock AIME 2024-202588.9%88.9%
ProofBench21%17%
LMArena Math14191455
FrontierMath (Feb 2025 set)28.3%39%
FrontierMath Tier 4 (v1)2.1%14.6%
FrontierMath (Tiers 1-3)51.2%—
FrontierMath Tier 49.8%—

Knowledge Muse Spark leads

GPT-5.4 mini: 51.5 (#67), Muse Spark: 65.7 (#13)

Knowledge benchmarks
BenchmarkGPT-5.4 miniMuse Spark
GPQA Diamond86.9%89.8%
LMArena Expert14351457
Humanity's Last Exam—40.6%
SimpleQA Verified29.4%—
Vectara Hallucination Rate5.5%—

Multimodal Muse Spark leads

GPT-5.4 mini: 39.7 (#56), Muse Spark: 43.4 (#24)

Multimodal benchmarks
BenchmarkGPT-5.4 miniMuse Spark
LMArena Vision12451306
LMArena Document—1444

Multilingual Muse Spark leads

GPT-5.4 mini: 51.9 (#96), Muse Spark: 56.1 (#24)

Multilingual benchmarks
BenchmarkGPT-5.4 miniMuse Spark
LMArena Non-English14051464
LMArena Chinese14461509
LMArena French14401497
LMArena German14091497
LMArena Korean13681459
LMArena Russian14171466
LMArena Spanish14051472
LMArena Japanese1374—

Instruction Following Muse Spark leads

GPT-5.4 mini: 74.1 (#102), Muse Spark: 75.9 (#51)

Instruction Following benchmarks
BenchmarkGPT-5.4 miniMuse Spark
LMArena Instruction Following14051442

Long Context Muse Spark leads

GPT-5.4 mini: 43.0 (#112), Muse Spark: 44.4 (#69)

Long Context benchmarks
BenchmarkGPT-5.4 miniMuse Spark
LMArena Longer Query14071451

Writing & Preference Muse Spark leads

GPT-5.4 mini: 64.0 (#58), Muse Spark: 66.0 (#39)

Writing & Preference benchmarks
BenchmarkGPT-5.4 miniMuse Spark
LMArena Text14121474
LMArena Creative Writing13701459
LMArena Multi-Turn14291477
EQ-Bench Creative Writing1665—

Frequently asked questions

Is GPT-5.4 mini better than Muse Spark?

Muse Spark is the stronger model overall, scoring 50.6 to 45.0 on the Noometry Index.

Is GPT-5.4 mini or Muse Spark better for coding?

They score almost the same on coding (45.2 vs 46.2); test both on your own repository before choosing.

How many benchmarks do GPT-5.4 mini and Muse Spark share?

25 benchmarks have published results for both models. GPT-5.4 mini has 46 scored results on Noometry and Muse Spark has 27.

Related comparisons

Go deeper