Model comparison

Mistral Nemo vs Muse Spark

Muse Spark is the stronger model overall, scoring 50.6 to 26.4 on the Noometry Index.

Last verified . 2 shared benchmarks.

Mistral Nemo Mistral AI

26.4

Rank #337 Confirmed

Muse Spark Meta

50.6

Rank #46 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Mistral Nemo scores higher in 0 categories and Muse Spark in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Muse Spark leads 65.7 to 12.3.
  • The biggest single-benchmark swing is GPQA Diamond: 29.9% for Mistral Nemo and 89.8% for Muse Spark.
  • Mistral Nemo has downloadable open weights; the other is API-only.

Side by side

Mistral Nemo and Muse Spark specifications
Mistral NemoMuse Spark
ProviderMistral AIMeta
Noometry Index26.450.6
Released2024-07-012026-04-08
WeightsOpenProprietary
Context window128K—
Max output128K—
Input $ / M tokens$0.15—
Output $ / M tokens$0.15—
Results tracked1027

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Mistral Nemo: —, Muse Spark: 46.2 (#69)

Coding benchmarks
BenchmarkMistral NemoMuse Spark
SciCode—51.5%
LMArena Coding—1481

Agentic & Tool Use Not comparable

Mistral Nemo: 23.5 (#125), Muse Spark: —

Agentic & Tool Use benchmarks
BenchmarkMistral NemoMuse Spark
Berkeley Function Calling Leaderboard27.6%—
BALROG17.6%—

Reasoning Muse Spark leads

Mistral Nemo: 20.7 (#232), Muse Spark: 35.9 (#67)

Reasoning benchmarks
BenchmarkMistral NemoMuse Spark
Epoch Capabilities Index118.68152.04
CritPt—11.3%
LMArena Hard Prompts—1474
DTBench48.6%—
PIQA83.5%—

Math Muse Spark leads

Mistral Nemo: 25.5 (#268), Muse Spark: 47.8 (#66)

Math benchmarks
BenchmarkMistral NemoMuse Spark
OTIS Mock AIME 2024-2025—88.9%
ProofBench—17%
LMArena Math—1455
MATH Level 510.8%—
FrontierMath (Feb 2025 set)—39%
FrontierMath Tier 4 (v1)—14.6%
GSM8K84.2%—

Knowledge Muse Spark leads

Mistral Nemo: 12.3 (#298), Muse Spark: 65.7 (#13)

Knowledge benchmarks
BenchmarkMistral NemoMuse Spark
GPQA Diamond29.9%89.8%
Humanity's Last Exam—40.6%
LMArena Expert—1457
BoolQ82.5%—

Multimodal Not comparable

Mistral Nemo: —, Muse Spark: 43.4 (#24)

Multimodal benchmarks
BenchmarkMistral NemoMuse Spark
LMArena Vision—1306
LMArena Document—1444

Multilingual Not comparable

Mistral Nemo: —, Muse Spark: 56.1 (#24)

Multilingual benchmarks
BenchmarkMistral NemoMuse Spark
LMArena Non-English—1464
LMArena Chinese—1509
LMArena French—1497
LMArena German—1497
LMArena Korean—1459
LMArena Russian—1466
LMArena Spanish—1472

Instruction Following Not comparable

Mistral Nemo: —, Muse Spark: 75.9 (#51)

Instruction Following benchmarks
BenchmarkMistral NemoMuse Spark
LMArena Instruction Following—1442

Long Context Not comparable

Mistral Nemo: —, Muse Spark: 44.4 (#69)

Long Context benchmarks
BenchmarkMistral NemoMuse Spark
LMArena Longer Query—1451

Writing & Preference Muse Spark leads

Mistral Nemo: 28.5 (#296), Muse Spark: 66.0 (#39)

Writing & Preference benchmarks
BenchmarkMistral NemoMuse Spark
LMArena Text—1474
LMArena Creative Writing—1459
EQ-Bench Creative Writing881—
LMArena Multi-Turn—1477

Frequently asked questions

Is Mistral Nemo better than Muse Spark?

Muse Spark is the stronger model overall, scoring 50.6 to 26.4 on the Noometry Index.

How many benchmarks do Mistral Nemo and Muse Spark share?

2 benchmarks have published results for both models. Mistral Nemo has 10 scored results on Noometry and Muse Spark has 27.

Related comparisons

Go deeper