Skip to content
Bencher

Models · Meta (Meta Superintelligence Labs, Muse) · Out sinceReleased 8 Apr 2026

Muse Spark

Only from the makerNot on OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

Muse Spark is made by Meta (Meta Superintelligence Labs, Muse). We don't have enough test results yet to rank it.

Original Muse Spark (pre-1.1); not on OpenRouter.

Writing & creativity
—
Research & analysis
—
Coding
—
Price
Price unknown— / —
Price per 1M (blended)Blended / 1M
—
MemoryContext
—
Longest answerMax output
—
Test resultsResults
28 (16 independent16 indep.)
Out sinceReleased
8 Apr 2026
Made byVendor
Meta (Meta Superintelligence Labs, Muse)
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

You can't choose how long this one thinks.

No adjustable reasoning setting listed.

Read more

Links

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
BabyVision39.9%Defaulteffort not statedMaker's own figureVendor-reportedMeta ↗9 Jul 2026—'Muse Spark' comparison column on the Muse Spark 1.1 page.
CharXiv Reasoning (no tools)88.9%Defaulteffort not statedMaker's own figureVendor-reportedMeta ↗9 Jul 2026—'Muse Spark' comparison column on the Muse Spark 1.1 page.
DeepSWE10.0%Defaulteffort not statedMaker's own figureVendor-reportedMeta ↗9 Jul 2026—'Muse Spark' comparison column on the Muse Spark 1.1 page.
FrontierScience Research38.0%MaxContemplating modeMaker's own figureVendor-reportedMeta ↗8 Apr 2026—Integer as published.
Humanity's Last Exam58.0%MaxContemplating modeMaker's own figureVendor-reportedMeta ↗8 Apr 2026—Contemplating mode = multiple agents reasoning in parallel; post does not say whether tools were used. Integer as published.
Humanity's Last Exam40.6%DefaultMuse SparkIndependent testIndependentScale AI SEAL ↗1 Oct 2026—rank 4 (Scale rank accounts for CI); ±1.92; entry added 2026-04-08
Humanity's Last Exam (with tools)50.4%Defaulteffort not statedMaker's own figureVendor-reportedMeta ↗9 Jul 2026—With tools. 'Muse Spark' comparison column on the Muse Spark 1.1 page.
JobBench17.0%Defaulteffort not statedMaker's own figureVendor-reportedMeta ↗9 Jul 2026—'Muse Spark' comparison column on the Muse Spark 1.1 page.
LMArena Text - Coding category1530Defaultmuse-sparkIndependent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 22 (CI rank 7-48); 95% CI 1520-1540; 3930 votes
LMArena Text - Hard Prompts1506Defaultmuse-sparkIndependent testIndependentLMArena ↗30 Sep 2026—rank 20 (rank range 9-41); 95% CI 1499.3-1513.1; 9145 votes; style-controlled
LMArena Text - Multi-Turn1494Defaultmuse-sparkIndependent testIndependentLMArena ↗30 Sep 2026—rank 18 (rank range 6-55); 95% CI 1480.4-1506.6; 2284 votes; style-controlled
LMArena Text - Non-English1477Defaultmuse-sparkIndependent testIndependentLMArena ↗30 Sep 2026—rank 18 (rank range 6-33); 95% CI 1470.0-1484.8; 7451 votes; style-controlled
LMArena Text - Occupational: Business, Management & Finance1489Defaultmuse-sparkIndependent testIndependentLMArena ↗30 Sep 2026—rank 13 (rank range 3-41); 95% CI 1477.7-1500.8; 2866 votes; style-controlled
LMArena Text - Occupational: Legal & Government1501Defaultmuse-sparkIndependent testIndependentLMArena ↗30 Sep 2026—rank 12 (rank range 1-60); 95% CI 1481.8-1520.4; 1016 votes; style-controlled
LMArena Text (overall)1489Defaultmuse-sparkIndependent testIndependentLMArena ↗30 Sep 2026—text overall, style control on; rank 15 (CI rank 5-27); 95% CI 1483-1495; 14128 votes
MCP Atlas82.2%DefaultMuse SparkIndependent testIndependentScale AI SEAL ↗1 Oct 2026—rank 2 (Scale rank accounts for CI); ±2.3; entry added 2026-04-08
MCP Atlas82.2%Defaulteffort not statedMaker's own figureVendor-reportedMeta ↗9 Jul 2026—'Muse Spark' comparison column on the Muse Spark 1.1 page.
OSWorld-Verified53.3%Defaulteffort not statedMaker's own figureVendor-reportedMeta ↗9 Jul 2026—'Muse Spark' comparison column on the Muse Spark 1.1 page.
PRBench Finance (Scale)52.4%DefaultMuse SparkIndependent testIndependentScale AI (SEAL) ↗8 Apr 2026—Scale rank 5; ±0.06 CI
PRBench Legal (Scale)52.3%DefaultMuse SparkIndependent testIndependentScale AI (SEAL) ↗8 Apr 2026—Scale rank 3; ±0.06 CI
SWE Atlas - Codebase QnA24.2%DefaultMuse SparkIndependent testIndependentScale AI SEAL ↗1 Oct 2026—rank 13 (Scale rank accounts for CI); ±4.6; entry added 2026-04-08
SWE Atlas - Test Writing31.1%DefaultMuse SparkIndependent testIndependentScale AI SEAL ↗1 Oct 2026—rank 5 (Scale rank accounts for CI); ±5.76; entry added 2026-04-08
SWE-Bench Pro (private/commercial set)44.7%DefaultMuse Spark*Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 3 (Scale rank accounts for CI); ±6.05; entry added 2026-04-08; * = Scale footnote (see page)
SWE-Bench Pro (public, v1)55.0%DefaultMuse Spark*Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 3 (Scale rank accounts for CI); ±3.6; entry added 2026-04-08; * = Scale footnote (see page)
SWE-Bench Pro (public, v1)55.0%Defaulteffort not statedMaker's own figureVendor-reportedMeta ↗9 Jul 2026—'Muse Spark' comparison column on the Muse Spark 1.1 page.
Terminal-Bench 2.167.3%Defaulteffort not statedMaker's own figureVendor-reportedMeta ↗9 Jul 2026—'Muse Spark' comparison column on the Muse Spark 1.1 page.
Toolathlon-Verified49.4%Defaulteffort not statedMaker's own figureVendor-reportedMeta ↗9 Jul 2026—'Muse Spark' comparison column on the Muse Spark 1.1 page.
Vals TaxEval v277.7%Defaultvals id meta/muse_sparkIndependent testIndependentVals.ai ↗1 Sep 2026—vals id meta/muse_spark; rank 3/145; ±0.806 stderr; $0.001922/test