Model comparison

Gemini 3.1 Pro Preview vs Muse Spark

Gemini 3.1 Pro Preview is the stronger model overall, scoring 56.7 to 50.6 on the Noometry Index.

Last verified . 27 shared benchmarks.

Gemini 3.1 Pro Preview Google

56.7

Rank #23 Confirmed

Muse Spark Meta

50.6

Rank #46 Confirmed

Summary

  • They share 27 benchmarks with published results for both. Gemini 3.1 Pro Preview scores higher in 7 categories and Muse Spark in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemini 3.1 Pro Preview leads 71.7 to 35.9.
  • The biggest single-benchmark swing is ProofBench: 26% for Gemini 3.1 Pro Preview and 17% for Muse Spark.

Side by side

Gemini 3.1 Pro Preview and Muse Spark specifications
Gemini 3.1 Pro PreviewMuse Spark
ProviderGoogleMeta
Noometry Index56.750.6
Released2026-02-192026-04-08
WeightsProprietaryProprietary
Context window1.05M—
Max output66K—
Input $ / M tokens$2—
Output $ / M tokens$12—
Results tracked7127

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Muse Spark leads

Gemini 3.1 Pro Preview: 42.5 (#99), Muse Spark: 46.2 (#69)

Coding benchmarks
BenchmarkGemini 3.1 Pro PreviewMuse Spark
SciCode58.9%51.5%
LMArena Coding14841481
SWE-bench Verified75.6%—
DeepSWE11.7%—
LMArena WebDev1447—
GSO22.6%—
WeirdML72.1%—
MirrorCode8.9%—
ALE-Bench1,161—
AlgoTune2.02—

Agentic & Tool Use Not comparable

Gemini 3.1 Pro Preview: 37.7 (#34), Muse Spark: —

Agentic & Tool Use benchmarks
BenchmarkGemini 3.1 Pro PreviewMuse Spark
Terminal-Bench80.2%—
APEX-Agents35.3%—
τ²-bench Banking26%—
DeepResearch Bench47.8%—
PostTrainBench22%—
BALROG57%—
ExploitBench26.1%—
GBAEval0.8%—
GDP.pdf17%—
LMArena Search1211—
METR Time Horizons77%—
Vending-Bench 23,774—

Reasoning Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 71.7 (#12), Muse Spark: 35.9 (#67)

Reasoning benchmarks
BenchmarkGemini 3.1 Pro PreviewMuse Spark
CritPt17.7%11.3%
LMArena Hard Prompts14851474
Epoch Capabilities Index154.77152.04
ARC-AGI-277.1%—
SimpleBench79.6%—
NYT Connections (extended)97.4%—
ARC-AGI-198%—
Chess Puzzles55%—
EnigmaEval36.8%—
Thematic Generalization79.4%—
EBR-Bench14.3%—
Mystery Game Puzzles34%—
DTBench97.1%—
LMCA53.8%—
ForecastBench59—

Math Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 62.1 (#34), Muse Spark: 47.8 (#66)

Math benchmarks
BenchmarkGemini 3.1 Pro PreviewMuse Spark
OTIS Mock AIME 2024-202595.6%88.9%
ProofBench26%17%
LMArena Math14851455
FrontierMath (Feb 2025 set)36.9%39%
FrontierMath Tier 4 (v1)16.7%14.6%
FrontierMath (Tiers 1-3)59.6%—
FrontierMath Tier 426.8%—
MathArena Final-Answer Competitions86.5%—

Knowledge Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 71.8 (#3), Muse Spark: 65.7 (#13)

Knowledge benchmarks
BenchmarkGemini 3.1 Pro PreviewMuse Spark
GPQA Diamond94.4%89.8%
Humanity's Last Exam46.4%40.6%
LMArena Expert14851457
SimpleQA Verified73.5%—
Vectara Hallucination Rate10.4%—

Multimodal Muse Spark leads

Gemini 3.1 Pro Preview: 37.9 (#69), Muse Spark: 43.4 (#24)

Multimodal benchmarks
BenchmarkGemini 3.1 Pro PreviewMuse Spark
LMArena Vision12961306
LMArena Document14441444
Blueprint-Bench 226.5%—
Furniture Assembly26.7%—

Multilingual Too close to call

Gemini 3.1 Pro Preview: 57.0 (#12), Muse Spark: 56.1 (#24)

Multilingual benchmarks
BenchmarkGemini 3.1 Pro PreviewMuse Spark
LMArena Non-English14771464
LMArena Chinese15291509
LMArena French14871497
LMArena German14911497
LMArena Korean14551459
LMArena Russian14981466
LMArena Spanish14791472
LMArena Japanese1493—

Instruction Following Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 77.0 (#32), Muse Spark: 75.9 (#51)

Instruction Following benchmarks
BenchmarkGemini 3.1 Pro PreviewMuse Spark
LMArena Instruction Following14661442

Long Context Gemini 3.1 Pro Preview leads

Gemini 3.1 Pro Preview: 47.4 (#18), Muse Spark: 44.4 (#69)

Long Context benchmarks
BenchmarkGemini 3.1 Pro PreviewMuse Spark
LMArena Longer Query14831451
CL-bench20.8%—
CL-bench Life16.9%—

Writing & Preference Too close to call

Gemini 3.1 Pro Preview: 66.1 (#37), Muse Spark: 66.0 (#39)

Writing & Preference benchmarks
BenchmarkGemini 3.1 Pro PreviewMuse Spark
LMArena Text14811474
LMArena Creative Writing14821459
LMArena Multi-Turn14881477
EQ-Bench Creative Writing1491—
EQ-Bench 41142—

Frequently asked questions

Is Gemini 3.1 Pro Preview better than Muse Spark?

Gemini 3.1 Pro Preview is the stronger model overall, scoring 56.7 to 50.6 on the Noometry Index.

Is Gemini 3.1 Pro Preview or Muse Spark better for coding?

Muse Spark scores higher on coding benchmarks: 46.2 versus 42.5 in the Noometry coding category.

How many benchmarks do Gemini 3.1 Pro Preview and Muse Spark share?

27 benchmarks have published results for both models. Gemini 3.1 Pro Preview has 71 scored results on Noometry and Muse Spark has 27.

Related comparisons

Go deeper