Model comparison

Muse Spark 1.3 vs Trinity Large Thinking

Muse Spark 1.3 is the stronger model overall, scoring 54.8 to 38.6 on the Noometry Index. Trinity Large Thinking costs 5.2× less per token, which makes it the better buy when Muse Spark 1.3's lead doesn't matter for your workload.

Last verified . 21 shared benchmarks.

Muse Spark 1.3 Meta

54.8

Rank #27 Confirmed

Trinity Large Thinking Arcee AI

38.6

Rank #185 Confirmed

Summary

  • They share 21 benchmarks with published results for both. Muse Spark 1.3 scores higher in 8 categories and Trinity Large Thinking in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Muse Spark 1.3 leads 54.0 to 16.9.
  • The biggest single-benchmark swing is NYT Connections (extended): 85.1% for Muse Spark 1.3 and 16.5% for Trinity Large Thinking.
  • Trinity Large Thinking is cheaper at $0.25 / $0.80 per million input/output tokens, against $1.25 / $4.25 for Muse Spark 1.3.
  • Muse Spark 1.3 accepts more context: 1.05M tokens versus 262K.
  • Trinity Large Thinking has downloadable open weights; the other is API-only.

Side by side

Muse Spark 1.3 and Trinity Large Thinking specifications
Muse Spark 1.3Trinity Large Thinking
ProviderMetaArcee AI
Noometry Index54.838.6
Released2026-09-022026-04-01
WeightsProprietaryOpen
Context window1.05M262K
Max output131K80K
Input $ / M tokens$1.25$0.25
Output $ / M tokens$4.25$0.80
Results tracked3724

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Muse Spark 1.3 leads

Muse Spark 1.3: 56.6 (#21), Trinity Large Thinking: 34.1 (#244)

Coding benchmarks
BenchmarkMuse Spark 1.3Trinity Large Thinking
LMArena WebDev16571238
SciCode59.7%36.1%
LMArena Coding15141381
CursorBench41.6%—

Agentic & Tool Use Not comparable

Muse Spark 1.3: 38.6 (#30), Trinity Large Thinking: —

Agentic & Tool Use benchmarks
BenchmarkMuse Spark 1.3Trinity Large Thinking
APEX-Agents57.8%—
GDP.pdf27.6%—

Reasoning Muse Spark 1.3 leads

Muse Spark 1.3: 54.0 (#27), Trinity Large Thinking: 16.9 (#298)

Reasoning benchmarks
BenchmarkMuse Spark 1.3Trinity Large Thinking
NYT Connections (extended)85.1%16.5%
CritPt26%0.9%
LMArena Hard Prompts15031350
Chess Puzzles38%—
Thematic Generalization—41.6%
Mystery Game Puzzles25%—
DTBench96.5%—
LMCA53.9%—
Surface Evolver Bench—15.6%
Bench to the Future 30.14—
Epoch Capabilities Index156.75—

Math Muse Spark 1.3 leads

Muse Spark 1.3: 73.1 (#21), Trinity Large Thinking: 37.6 (#149)

Math benchmarks
BenchmarkMuse Spark 1.3Trinity Large Thinking
LMArena Math14941366
FrontierMath (Tiers 1-3)74.4%—
FrontierMath Tier 446.3%—
OTIS Mock AIME 2024-202599.2%—
ProofBench58%—

Knowledge Muse Spark 1.3 leads

Muse Spark 1.3: 42.6 (#95), Trinity Large Thinking: 40.9 (#113)

Knowledge benchmarks
BenchmarkMuse Spark 1.3Trinity Large Thinking
LMArena Expert15161360
Vectara Hallucination Rate—6.9%

Multimodal Not comparable

Muse Spark 1.3: 43.7 (#22), Trinity Large Thinking: —

Multimodal benchmarks
BenchmarkMuse Spark 1.3Trinity Large Thinking
LMArena Vision1309—
LMArena Document1471—

Multilingual Muse Spark 1.3 leads

Muse Spark 1.3: 57.4 (#8), Trinity Large Thinking: 46.2 (#160)

Multilingual benchmarks
BenchmarkMuse Spark 1.3Trinity Large Thinking
LMArena Non-English14811325
LMArena Chinese15291373
LMArena French15241374
LMArena German15151356
LMArena Japanese14741311
LMArena Korean15011306
LMArena Russian14901337
LMArena Spanish14901357

Instruction Following Muse Spark 1.3 leads

Muse Spark 1.3: 77.5 (#22), Trinity Large Thinking: 70.5 (#162)

Instruction Following benchmarks
BenchmarkMuse Spark 1.3Trinity Large Thinking
LMArena Instruction Following14771334

Long Context Muse Spark 1.3 leads

Muse Spark 1.3: 45.6 (#32), Trinity Large Thinking: 41.3 (#144)

Long Context benchmarks
BenchmarkMuse Spark 1.3Trinity Large Thinking
LMArena Longer Query14881355

Writing & Preference Muse Spark 1.3 leads

Muse Spark 1.3: 73.6 (#9), Trinity Large Thinking: 53.8 (#158)

Writing & Preference benchmarks
BenchmarkMuse Spark 1.3Trinity Large Thinking
LMArena Text14901340
LMArena Creative Writing14551320
LMArena Multi-Turn14821342
EQ-Bench Creative Writing1906—

Frequently asked questions

Is Muse Spark 1.3 better than Trinity Large Thinking?

Muse Spark 1.3 is the stronger model overall, scoring 54.8 to 38.6 on the Noometry Index. Trinity Large Thinking costs 5.2× less per token, which makes it the better buy when Muse Spark 1.3's lead doesn't matter for your workload.

Which is cheaper, Muse Spark 1.3 or Trinity Large Thinking?

Trinity Large Thinking is cheaper. It lists at $0.25 per million input tokens and $0.80 per million output tokens; Muse Spark 1.3 lists at $1.25 and $4.25.

Is Muse Spark 1.3 or Trinity Large Thinking better for coding?

Muse Spark 1.3 scores higher on coding benchmarks: 56.6 versus 34.1 in the Noometry coding category.

Which has the bigger context window?

Muse Spark 1.3 does, with 1.05M tokens against 262K.

How many benchmarks do Muse Spark 1.3 and Trinity Large Thinking share?

21 benchmarks have published results for both models. Muse Spark 1.3 has 37 scored results on Noometry and Trinity Large Thinking has 24.

Related comparisons

Go deeper