Model comparison

ERNIE 5.1 vs Muse Spark 1.2

Muse Spark 1.2 is the stronger model overall, scoring 50.3 to 43.8 on the Noometry Index.

Last verified . 15 shared benchmarks.

ERNIE 5.1 Baidu

43.8

Rank #83 Confirmed

Muse Spark 1.2 Meta

50.3

Rank #48 Confirmed

Summary

  • They share 15 benchmarks with published results for both. ERNIE 5.1 scores higher in 0 categories and Muse Spark 1.2 in 8 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Muse Spark 1.2 leads 51.3 to 21.9.
  • The biggest single-benchmark swing is NYT Connections (extended): 23.4% for ERNIE 5.1 and 79.2% for Muse Spark 1.2.

Side by side

ERNIE 5.1 and Muse Spark 1.2 specifications
ERNIE 5.1Muse Spark 1.2
ProviderBaiduMeta
Noometry Index43.850.3
Released—2026-08-05
WeightsProprietaryProprietary
Context window—1.05M
Max output—131K
Input $ / M tokens—$1.25
Output $ / M tokens—$4.25
Results tracked1931

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Muse Spark 1.2 leads

ERNIE 5.1: 44.0 (#76), Muse Spark 1.2: 49.2 (#51)

Coding benchmarks
BenchmarkERNIE 5.1Muse Spark 1.2
LMArena Coding14881495
DeepSWE—54.9%
LMArena WebDev—1533
FrontierSWE—12%
SciCode—56.4%
WeirdML—60.3%

Agentic & Tool Use Not comparable

ERNIE 5.1: —, Muse Spark 1.2: 29.4 (#87)

Agentic & Tool Use benchmarks
BenchmarkERNIE 5.1Muse Spark 1.2
APEX-Agents—36.4%
GDP.pdf—16%
LMArena Search1227—

Reasoning Muse Spark 1.2 leads

ERNIE 5.1: 21.9 (#211), Muse Spark 1.2: 51.3 (#34)

Reasoning benchmarks
BenchmarkERNIE 5.1Muse Spark 1.2
NYT Connections (extended)23.4%79.2%
LMArena Hard Prompts14811486
SimpleBench—74.5%
CritPt—17.7%
DTBench—94.7%
LMCA—48.4%
Epoch Capabilities Index—154.87

Math Muse Spark 1.2 leads

ERNIE 5.1: 40.3 (#92), Muse Spark 1.2: 46.4 (#70)

Math benchmarks
BenchmarkERNIE 5.1Muse Spark 1.2
LMArena Math14811471
ProofBench—43%

Knowledge Muse Spark 1.2 leads

ERNIE 5.1: 41.9 (#102), Muse Spark 1.2: 54.1 (#53)

Knowledge benchmarks
BenchmarkERNIE 5.1Muse Spark 1.2
LMArena Expert14921480
SimpleQA Verified—60.3%

Multimodal Not comparable

ERNIE 5.1: —, Muse Spark 1.2: 43.4 (#25)

Multimodal benchmarks
BenchmarkERNIE 5.1Muse Spark 1.2
LMArena Vision—1305

Multilingual Muse Spark 1.2 leads

ERNIE 5.1: 55.5 (#29), Muse Spark 1.2: 57.1 (#11)

Multilingual benchmarks
BenchmarkERNIE 5.1Muse Spark 1.2
LMArena Non-English14541478
LMArena Chinese15081511
LMArena French14881513
LMArena Russian14591487
LMArena Spanish14731498
LMArena German1470—
LMArena Japanese1422—
LMArena Korean1427—

Instruction Following Too close to call

ERNIE 5.1: 76.7 (#37), Muse Spark 1.2: 76.7 (#36)

Instruction Following benchmarks
BenchmarkERNIE 5.1Muse Spark 1.2
LMArena Instruction Following14601461

Long Context Too close to call

ERNIE 5.1: 44.7 (#59), Muse Spark 1.2: 45.2 (#48)

Long Context benchmarks
BenchmarkERNIE 5.1Muse Spark 1.2
LMArena Longer Query14621475

Writing & Preference Muse Spark 1.2 leads

ERNIE 5.1: 65.1 (#52), Muse Spark 1.2: 72.3 (#14)

Writing & Preference benchmarks
BenchmarkERNIE 5.1Muse Spark 1.2
LMArena Text14681482
LMArena Creative Writing14411449
LMArena Multi-Turn14711494
EQ-Bench Creative Writing—1840

Frequently asked questions

Is ERNIE 5.1 better than Muse Spark 1.2?

Muse Spark 1.2 is the stronger model overall, scoring 50.3 to 43.8 on the Noometry Index.

Is ERNIE 5.1 or Muse Spark 1.2 better for coding?

Muse Spark 1.2 scores higher on coding benchmarks: 49.2 versus 44.0 in the Noometry coding category.

How many benchmarks do ERNIE 5.1 and Muse Spark 1.2 share?

15 benchmarks have published results for both models. ERNIE 5.1 has 19 scored results on Noometry and Muse Spark 1.2 has 31.

Related comparisons

Go deeper