Skip to content
Bencher

Models · Meta (Meta Superintelligence Labs, Muse) · Out sinceReleased 9 Jul 2026

Muse Spark 1.1#11 for writing.#11 for writing, best at its default setting.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

Muse Spark 1.1 is made by Meta (Meta Superintelligence Labs, Muse). Among the models we track it ranks #11 for writing, #18 for research and analysis, #18 for coding. It's mid-priced to use.

Launched with the Meta Model API public preview. Price from OpenRouter listing (same as 1.2/1.3). Evaluation report link currently returns an error.

Writing & creativity
58.4 / 100 · #11
Research & analysis
65.4 / 100 · #18
Coding
51.1 / 100 · #18
Price
Mid-priced$1.25 / $4.25
Price per 1M (blended)Blended / 1M
$2
MemoryContext
1M
Longest answerMax output
—
Test resultsResults
45 (31 independent31 indep.)
Out sinceReleased
9 Jul 2026
Made byVendor
Meta (Meta Superintelligence Labs, Muse)
UnderstandsInputs
text, image, audio, video

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as meta/muse-spark-1.1.

Route it as meta/muse-spark-1.1 at $1.25 in / $4.25 out per 1M tokens, 1M context. Listed since 16 Jul 2026.

include_reasoningmax_tokensreasoningreasoning_effortrepetition_penaltyresponse_formatstructured_outputstemperaturetool_choicetoolstop_ktop_p

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
Artificial Analysis Intelligence Index33.7Extra highMuse Spark 1.1 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug muse-spark-1-1; list price $1.25/4.25 per 1M in/out; cost to run AA Intelligence Index $1.38/task (AA marks this variant deprecated)
BabyVision76.3%Defaulteffort not statedMaker's own figureVendor-reportedMeta ↗9 Jul 2026—Effort not stated on page (Meta's later comparisons run 1.1 at xhigh). Evaluation report link currently broken.
CharXiv Reasoning (no tools)88.4%Defaulteffort not statedMaker's own figureVendor-reportedMeta ↗9 Jul 2026—Effort not stated on page (Meta's later comparisons run 1.1 at xhigh). Evaluation report link currently broken.
DeepSWE53.0%Extra highMaker's own figureVendor-reportedMeta ↗5 Aug 2026mini-swe-agentComparison bar in the Muse Spark 1.2 post.
DeepSWE53.3%Defaulteffort not statedMaker's own figureVendor-reportedMeta ↗9 Jul 2026—Effort not stated on page (Meta's later comparisons run 1.1 at xhigh). Evaluation report link currently broken.
EQ-Bench Creative Writing v3 (Elo)1927Defaultmuse-spark-1.1Independent testIndependentEQ-Bench ↗1 Oct 2026—leaderboard rank 13; rubric score 16.54/20; slop 12.11; avg length 7551 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench
GDPval-AA v21371Extra highMaker's own figureVendor-reportedMeta ↗5 Aug 2026AA StirrupComparison bar in the Muse Spark 1.2 post.
GPQA Diamond91.2%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗1 Sep 2026—±2.104 stderr; $0.034/test
GPQA Diamond89.8%Extra highMuse Spark 1.1 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug muse-spark-1-1) (AA marks this variant deprecated)
Humanity's Last Exam46.2%Extra highMuse Spark 1.1 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug muse-spark-1-1) (AA marks this variant deprecated)
Humanity's Last Exam (with tools)62.1%Defaulteffort not statedMaker's own figureVendor-reportedMeta ↗9 Jul 2026—With tools. Effort not stated on page (Meta's later comparisons run 1.1 at xhigh). Evaluation report link currently broken.
JobBench54.7%Defaulteffort not statedMaker's own figureVendor-reportedMeta ↗9 Jul 2026—Effort not stated on page (Meta's later comparisons run 1.1 at xhigh). Evaluation report link currently broken.
LLM Creative Story-Writing Benchmark (Lech Mazur)0.8HighMuse Spark 1.1 (high)Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 19/56; Thurstone comparison score (centered at 0); est. win chance 62%; 95% bootstrap 0.672 to 0.876
LMArena Text - Coding category1534Defaultmuse-spark-1.1Independent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 13 (CI rank 6-31); 95% CI 1527-1541; 10381 votes
LMArena Text - Hard Prompts1511Defaultmuse-spark-1.1Independent testIndependentLMArena ↗30 Sep 2026—rank 15 (rank range 7-30); 95% CI 1506.2-1516.5; 24296 votes; style-controlled
LMArena Text - Multi-Turn1496Defaultmuse-spark-1.1Independent testIndependentLMArena ↗30 Sep 2026—rank 13 (rank range 7-40); 95% CI 1487.4-1504.5; 5616 votes; style-controlled
LMArena Text - Non-English1484Defaultmuse-spark-1.1Independent testIndependentLMArena ↗30 Sep 2026—rank 11 (rank range 2-21); 95% CI 1478.8-1489.3; 21724 votes; style-controlled
LMArena Text - Occupational: Business, Management & Finance1494Defaultmuse-spark-1.1Independent testIndependentLMArena ↗30 Sep 2026—rank 9 (rank range 3-29); 95% CI 1485.8-1501.4; 7061 votes; style-controlled
LMArena Text - Occupational: Legal & Government1498Defaultmuse-spark-1.1Independent testIndependentLMArena ↗30 Sep 2026—rank 14 (rank range 1-50); 95% CI 1486.7-1509.2; 3156 votes; style-controlled
LMArena Text (overall)1492Defaultmuse-spark-1.1Independent testIndependentLMArena ↗30 Sep 2026—text overall, style control on; rank 12 (CI rank 5-22); 95% CI 1487-1496; 36558 votes
MCP Atlas88.1%Extra highMaker's own figureVendor-reportedMeta ↗5 Aug 2026—Comparison bar in the Muse Spark 1.2 post; same value on the 1.1 model page.
MCP Atlas88.1%DefaultMuse Spark 1.1Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 1 (Scale rank accounts for CI); ±1.95; entry added 2026-07-09
OSWorld-Verified80.8%Defaulteffort not statedMaker's own figureVendor-reportedMeta ↗9 Jul 2026—Effort not stated on page (Meta's later comparisons run 1.1 at xhigh). Evaluation report link currently broken.
PRBench Finance (Scale)55.0%DefaultMuse Spark 1.1Independent testIndependentScale AI (SEAL) ↗9 Jul 2026—Scale rank 2; ±0.14 CI
PRBench Legal (Scale)57.0%DefaultMuse Spark 1.1Independent testIndependentScale AI (SEAL) ↗9 Jul 2026—Scale rank 2; ±0.22 CI
ProgramBench (avg test pass rate)47.0%Extra highIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentavg behavioral-test pass rate ("Score" column); fully resolved 0.0%, almost (>=95% tests) 4.0%; avg cost $0.73/task
ProgramBench (fully resolved)0.0%Extra highIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentstrict fully-resolved rate; almost (>=95% tests) 4.0%; avg cost $0.73/task
SciCode58.8%Extra highMuse Spark 1.1 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug muse-spark-1-1) (AA marks this variant deprecated)
SimpleQA Verified57.8%Defaultmuse-spark-1.1Independent testIndependentEpoch AI ↗31 Aug 2026—Epoch-run (no tools); ±1.57 stderr
SkillsBench59.2%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗27 Sep 2026OpenHands±4.444 stderr; $0.449/test
SWE Atlas - Codebase QnA42.2%Extra highMuse Spark 1.1 (Mini-SWE-Agent) xHighIndependent testIndependentScale AI SEAL ↗1 Oct 2026mini-SWE-agentrank 2 (Scale rank accounts for CI); ±5.08; entry added 2026-07-09
SWE Atlas - Test Writing41.5%Extra highMuse Spark 1.1 (Mini-SWE-Agent) xHighIndependent testIndependentScale AI SEAL ↗1 Oct 2026mini-SWE-agentrank 2 (Scale rank accounts for CI); ±5.96; entry added 2026-07-09
SWE-Bench Pro (private/commercial set)51.5%DefaultMuse Spark 1.1*Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 1 (Scale rank accounts for CI); ±5.5; entry added 2026-07-09; * = Scale footnote (see page)
SWE-Bench Pro (public, v1)61.5%DefaultMuse Spark 1.1*Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 1 (Scale rank accounts for CI); ±3.1; entry added 2026-07-09; * = Scale footnote (see page)
SWE-Bench Pro (public, v1)61.5%Defaulteffort not statedMaker's own figureVendor-reportedMeta ↗9 Jul 2026—Effort not stated on page (Meta's later comparisons run 1.1 at xhigh). Evaluation report link currently broken.
SWE-bench Verified82.0%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗1 Sep 2026bash-only (single bash tool) agent±1.72 stderr; $0.352/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated)
Terminal-Bench 2.177.9%Extra highMuse Spark 1.1 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug muse-spark-1-1) (AA marks this variant deprecated)
Terminal-Bench 2.176.2%Extra highMaker's own figureVendor-reportedMeta ↗5 Aug 2026mini-swe-agentComparison bar in the Muse Spark 1.2 post. Avg of 5 attempts, pass@1. Methodology: https://research.meta.ai/static/muse-spark-1-2-methodology
Terminal-Bench 2.176.2%DefaultIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026mini-SWE-agentTB 2.1 (89 tasks, archived); effort not listed
Terminal-Bench 2.180.0%Defaulteffort not statedMaker's own figureVendor-reportedMeta ↗9 Jul 2026—Harness not stated; Meta's 1.2 post shows 76.2% with mini-swe-agent. Effort not stated on page (Meta's later comparisons run 1.1 at xhigh). Evaluation report link currently broken.
Terminal-Bench 4.06.1%Extra highMuse Spark 1.1 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug muse-spark-1-1) (AA marks this variant deprecated)
Toolathlon-Verified75.6%Defaulteffort not statedMaker's own figureVendor-reportedMeta ↗9 Jul 2026—Effort not stated on page (Meta's later comparisons run 1.1 at xhigh). Evaluation report link currently broken.
Vals CorpFin v271.3%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗12 Aug 2026—vals id meta/muse_spark_1_1; rank 4/134; ±0.889 stderr; $0.10923/test
Vals Finance Agent v257.2%Defaulteffort not statedMaker's own figureVendor-reportedMeta ↗9 Jul 2026—Effort not stated on page (Meta's later comparisons run 1.1 at xhigh). Evaluation report link currently broken.
Vals TaxEval v279.7%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗1 Sep 2026—vals id meta/muse_spark_1_1; rank 2/145; ±0.77 stderr; $0.014921/test