Model comparison

GPT-6 Astra vs Muse Spark 1.3

GPT-6 Astra is the stronger model overall, scoring 70.8 to 54.8 on the Noometry Index. Muse Spark 1.3 costs 10× less per token, which makes it the better buy when GPT-6 Astra's lead doesn't matter for your workload.

Last verified . 36 shared benchmarks.

GPT-6 Astra OpenAI

70.8

Rank #1 Confirmed

Muse Spark 1.3 Meta

54.8

Rank #27 Confirmed

Summary

  • They share 36 benchmarks with published results for both. GPT-6 Astra scores higher in 7 categories and Muse Spark 1.3 in 3 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where GPT-6 Astra leads 75.3 to 42.6.
  • The biggest single-benchmark swing is Mystery Game Puzzles: 84% for GPT-6 Astra and 25% for Muse Spark 1.3.
  • Muse Spark 1.3 is cheaper at $1.25 / $4.25 per million input/output tokens, against $10 / $50 for GPT-6 Astra.
  • GPT-6 Astra accepts more context: 1.05M tokens versus 1.05M.

Side by side

GPT-6 Astra and Muse Spark 1.3 specifications
GPT-6 AstraMuse Spark 1.3
ProviderOpenAIMeta
Noometry Index70.854.8
Released2026-09-032026-09-02
WeightsProprietaryProprietary
Context window1.05M1.05M
Max output128K131K
Input $ / M tokens$10$1.25
Output $ / M tokens$50$4.25
Results tracked5637

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-6 Astra leads

GPT-6 Astra: 73.7 (#2), Muse Spark 1.3: 56.6 (#21)

Coding benchmarks
BenchmarkGPT-6 AstraMuse Spark 1.3
LMArena WebDev17861657
SciCode56.5%59.7%
LMArena Coding14871514
DeepSWE74.1%—
FrontierCode53.3%—
CursorBench—41.6%
FrontierSWE65.5%—
GSO79.4%—
WeirdML93.6%—
MirrorCode46.7%—
ALE-Bench2,951—

Agentic & Tool Use GPT-6 Astra leads

GPT-6 Astra: 52.9 (#3), Muse Spark 1.3: 38.6 (#30)

Agentic & Tool Use benchmarks
BenchmarkGPT-6 AstraMuse Spark 1.3
APEX-Agents64.7%57.8%
GDP.pdf34.2%27.6%
Remote Labor Index20.8%—
BALROG68.3%—
Vending-Bench 215,515—

Reasoning GPT-6 Astra leads

GPT-6 Astra: 85.1 (#1), Muse Spark 1.3: 54.0 (#27)

Reasoning benchmarks
BenchmarkGPT-6 AstraMuse Spark 1.3
NYT Connections (extended)98.1%85.1%
CritPt31.7%26%
Chess Puzzles72%38%
LMArena Hard Prompts14621503
Mystery Game Puzzles84%25%
DTBench97.3%96.5%
LMCA64.4%53.9%
Bench to the Future 30.140.14
Epoch Capabilities Index166.45156.75
ARC-AGI-295%—
ARC-AGI-198.5%—
EBR-Bench76.2%—

Math GPT-6 Astra leads

GPT-6 Astra: 93.5 (#2), Muse Spark 1.3: 73.1 (#21)

Math benchmarks
BenchmarkGPT-6 AstraMuse Spark 1.3
FrontierMath (Tiers 1-3)93.7%74.4%
FrontierMath Tier 497.6%46.3%
OTIS Mock AIME 2024-2025100%99.2%
ProofBench99%58%
LMArena Math14651494
FrontierMath Erdős2.9%—

Knowledge GPT-6 Astra leads

GPT-6 Astra: 75.3 (#1), Muse Spark 1.3: 42.6 (#95)

Knowledge benchmarks
BenchmarkGPT-6 AstraMuse Spark 1.3
LMArena Expert14831516
GPQA Diamond95.8%—
Humanity's Last Exam54.8%—
SimpleQA Verified75.6%—
Vectara Hallucination Rate8.7%—

Multimodal GPT-6 Astra leads

GPT-6 Astra: 55.0 (#3), Muse Spark 1.3: 43.7 (#22)

Multimodal benchmarks
BenchmarkGPT-6 AstraMuse Spark 1.3
LMArena Vision12811309
LMArena Document14681471
Blueprint-Bench 249.7%—
Furniture Assembly80%—

Multilingual Muse Spark 1.3 leads

GPT-6 Astra: 53.7 (#61), Muse Spark 1.3: 57.4 (#8)

Multilingual benchmarks
BenchmarkGPT-6 AstraMuse Spark 1.3
LMArena Non-English14301481
LMArena Chinese14841529
LMArena French14561524
LMArena German14401515
LMArena Japanese13791474
LMArena Korean14261501
LMArena Russian14361490
LMArena Spanish14071490

Instruction Following Muse Spark 1.3 leads

GPT-6 Astra: 76.3 (#44), Muse Spark 1.3: 77.5 (#22)

Instruction Following benchmarks
BenchmarkGPT-6 AstraMuse Spark 1.3
LMArena Instruction Following14501477

Long Context Muse Spark 1.3 leads

GPT-6 Astra: 44.5 (#62), Muse Spark 1.3: 45.6 (#32)

Long Context benchmarks
BenchmarkGPT-6 AstraMuse Spark 1.3
LMArena Longer Query14561488

Writing & Preference GPT-6 Astra leads

GPT-6 Astra: 75.3 (#7), Muse Spark 1.3: 73.6 (#9)

Writing & Preference benchmarks
BenchmarkGPT-6 AstraMuse Spark 1.3
LMArena Text14411490
LMArena Creative Writing14181455
EQ-Bench Creative Writing21731906
LMArena Multi-Turn14481482

Frequently asked questions

Is GPT-6 Astra better than Muse Spark 1.3?

GPT-6 Astra is the stronger model overall, scoring 70.8 to 54.8 on the Noometry Index. Muse Spark 1.3 costs 10× less per token, which makes it the better buy when GPT-6 Astra's lead doesn't matter for your workload.

Which is cheaper, GPT-6 Astra or Muse Spark 1.3?

Muse Spark 1.3 is cheaper. It lists at $1.25 per million input tokens and $4.25 per million output tokens; GPT-6 Astra lists at $10 and $50.

Is GPT-6 Astra or Muse Spark 1.3 better for coding?

GPT-6 Astra scores higher on coding benchmarks: 73.7 versus 56.6 in the Noometry coding category.

Which has the bigger context window?

GPT-6 Astra does, with 1.05M tokens against 1.05M.

How many benchmarks do GPT-6 Astra and Muse Spark 1.3 share?

36 benchmarks have published results for both models. GPT-6 Astra has 56 scored results on Noometry and Muse Spark 1.3 has 37.

Related comparisons

Go deeper