Models · Meta (Meta Superintelligence Labs, Muse) · Out sinceReleased 9 Jul 2026
Muse Spark 1.1#11 for writing.#11 for writing, best at its default setting.
Muse Spark 1.1 is made by Meta (Meta Superintelligence Labs, Muse). Among the models we track it ranks #11 for writing, #18 for research and analysis, #18 for coding. It's mid-priced to use.
Launched with the Meta Model API public preview. Price from OpenRouter listing (same as 1.2/1.3). Evaluation report link currently returns an error.
- 58.4 / 100 · #11
- 65.4 / 100 · #18
- 51.1 / 100 · #18
- Mid-priced$1.25 / $4.25
- $2
- 1M
- —
- 45 (31 independent31 indep.)
- 9 Jul 2026
- Meta (Meta Superintelligence Labs, Muse)
- text, image, audio, video
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
You can't choose how long this one thinks.
No adjustable reasoning setting listed.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as meta/muse-spark-1.1.
Route it as meta/muse-spark-1.1 at $1.25 in / $4.25 out per 1M tokens, 1M context. Listed since 16 Jul 2026.
include_reasoningmax_tokensreasoningreasoning_effortrepetition_penaltyresponse_formatstructured_outputstemperaturetool_choicetoolstop_ktop_p
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| Artificial Analysis Intelligence Index | 33.7 | Extra highMuse Spark 1.1 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug muse-spark-1-1; list price $1.25/4.25 per 1M in/out; cost to run AA Intelligence Index $1.38/task (AA marks this variant deprecated) |
| BabyVision | 76.3% | Defaulteffort not stated | Maker's own figureVendor-reportedMeta ↗ | 9 Jul 2026 | — | Effort not stated on page (Meta's later comparisons run 1.1 at xhigh). Evaluation report link currently broken. |
| CharXiv Reasoning (no tools) | 88.4% | Defaulteffort not stated | Maker's own figureVendor-reportedMeta ↗ | 9 Jul 2026 | — | Effort not stated on page (Meta's later comparisons run 1.1 at xhigh). Evaluation report link currently broken. |
| DeepSWE | 53.0% | Extra high | Maker's own figureVendor-reportedMeta ↗ | 5 Aug 2026 | mini-swe-agent | Comparison bar in the Muse Spark 1.2 post. |
| DeepSWE | 53.3% | Defaulteffort not stated | Maker's own figureVendor-reportedMeta ↗ | 9 Jul 2026 | — | Effort not stated on page (Meta's later comparisons run 1.1 at xhigh). Evaluation report link currently broken. |
| EQ-Bench Creative Writing v3 (Elo) | 1927 | Defaultmuse-spark-1.1 | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 13; rubric score 16.54/20; slop 12.11; avg length 7551 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench |
| GDPval-AA v2 | 1371 | Extra high | Maker's own figureVendor-reportedMeta ↗ | 5 Aug 2026 | AA Stirrup | Comparison bar in the Muse Spark 1.2 post. |
| GPQA Diamond | 91.2% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±2.104 stderr; $0.034/test |
| GPQA Diamond | 89.8% | Extra highMuse Spark 1.1 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug muse-spark-1-1) (AA marks this variant deprecated) |
| Humanity's Last Exam | 46.2% | Extra highMuse Spark 1.1 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug muse-spark-1-1) (AA marks this variant deprecated) |
| Humanity's Last Exam (with tools) | 62.1% | Defaulteffort not stated | Maker's own figureVendor-reportedMeta ↗ | 9 Jul 2026 | — | With tools. Effort not stated on page (Meta's later comparisons run 1.1 at xhigh). Evaluation report link currently broken. |
| JobBench | 54.7% | Defaulteffort not stated | Maker's own figureVendor-reportedMeta ↗ | 9 Jul 2026 | — | Effort not stated on page (Meta's later comparisons run 1.1 at xhigh). Evaluation report link currently broken. |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | 0.8 | HighMuse Spark 1.1 (high) | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 19/56; Thurstone comparison score (centered at 0); est. win chance 62%; 95% bootstrap 0.672 to 0.876 |
| LMArena Text - Coding category | 1534 | Defaultmuse-spark-1.1 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text arena coding category, style control on; rank 13 (CI rank 6-31); 95% CI 1527-1541; 10381 votes |
| LMArena Text - Hard Prompts | 1511 | Defaultmuse-spark-1.1 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 15 (rank range 7-30); 95% CI 1506.2-1516.5; 24296 votes; style-controlled |
| LMArena Text - Multi-Turn | 1496 | Defaultmuse-spark-1.1 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 13 (rank range 7-40); 95% CI 1487.4-1504.5; 5616 votes; style-controlled |
| LMArena Text - Non-English | 1484 | Defaultmuse-spark-1.1 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 11 (rank range 2-21); 95% CI 1478.8-1489.3; 21724 votes; style-controlled |
| LMArena Text - Occupational: Business, Management & Finance | 1494 | Defaultmuse-spark-1.1 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 9 (rank range 3-29); 95% CI 1485.8-1501.4; 7061 votes; style-controlled |
| LMArena Text - Occupational: Legal & Government | 1498 | Defaultmuse-spark-1.1 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 14 (rank range 1-50); 95% CI 1486.7-1509.2; 3156 votes; style-controlled |
| LMArena Text (overall) | 1492 | Defaultmuse-spark-1.1 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text overall, style control on; rank 12 (CI rank 5-22); 95% CI 1487-1496; 36558 votes |
| MCP Atlas | 88.1% | Extra high | Maker's own figureVendor-reportedMeta ↗ | 5 Aug 2026 | — | Comparison bar in the Muse Spark 1.2 post; same value on the 1.1 model page. |
| MCP Atlas | 88.1% | DefaultMuse Spark 1.1 | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 1 (Scale rank accounts for CI); ±1.95; entry added 2026-07-09 |
| OSWorld-Verified | 80.8% | Defaulteffort not stated | Maker's own figureVendor-reportedMeta ↗ | 9 Jul 2026 | — | Effort not stated on page (Meta's later comparisons run 1.1 at xhigh). Evaluation report link currently broken. |
| PRBench Finance (Scale) | 55.0% | DefaultMuse Spark 1.1 | Independent testIndependentScale AI (SEAL) ↗ | 9 Jul 2026 | — | Scale rank 2; ±0.14 CI |
| PRBench Legal (Scale) | 57.0% | DefaultMuse Spark 1.1 | Independent testIndependentScale AI (SEAL) ↗ | 9 Jul 2026 | — | Scale rank 2; ±0.22 CI |
| ProgramBench (avg test pass rate) | 47.0% | Extra high | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | avg behavioral-test pass rate ("Score" column); fully resolved 0.0%, almost (>=95% tests) 4.0%; avg cost $0.73/task |
| ProgramBench (fully resolved) | 0.0% | Extra high | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | strict fully-resolved rate; almost (>=95% tests) 4.0%; avg cost $0.73/task |
| SciCode | 58.8% | Extra highMuse Spark 1.1 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug muse-spark-1-1) (AA marks this variant deprecated) |
| SimpleQA Verified | 57.8% | Defaultmuse-spark-1.1 | Independent testIndependentEpoch AI ↗ | 31 Aug 2026 | — | Epoch-run (no tools); ±1.57 stderr |
| SkillsBench | 59.2% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | OpenHands | ±4.444 stderr; $0.449/test |
| SWE Atlas - Codebase QnA | 42.2% | Extra highMuse Spark 1.1 (Mini-SWE-Agent) xHigh | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | mini-SWE-agent | rank 2 (Scale rank accounts for CI); ±5.08; entry added 2026-07-09 |
| SWE Atlas - Test Writing | 41.5% | Extra highMuse Spark 1.1 (Mini-SWE-Agent) xHigh | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | mini-SWE-agent | rank 2 (Scale rank accounts for CI); ±5.96; entry added 2026-07-09 |
| SWE-Bench Pro (private/commercial set) | 51.5% | DefaultMuse Spark 1.1* | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 1 (Scale rank accounts for CI); ±5.5; entry added 2026-07-09; * = Scale footnote (see page) |
| SWE-Bench Pro (public, v1) | 61.5% | DefaultMuse Spark 1.1* | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 1 (Scale rank accounts for CI); ±3.1; entry added 2026-07-09; * = Scale footnote (see page) |
| SWE-Bench Pro (public, v1) | 61.5% | Defaulteffort not stated | Maker's own figureVendor-reportedMeta ↗ | 9 Jul 2026 | — | Effort not stated on page (Meta's later comparisons run 1.1 at xhigh). Evaluation report link currently broken. |
| SWE-bench Verified | 82.0% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | bash-only (single bash tool) agent | ±1.72 stderr; $0.352/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated) |
| Terminal-Bench 2.1 | 77.9% | Extra highMuse Spark 1.1 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug muse-spark-1-1) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 76.2% | Extra high | Maker's own figureVendor-reportedMeta ↗ | 5 Aug 2026 | mini-swe-agent | Comparison bar in the Muse Spark 1.2 post. Avg of 5 attempts, pass@1. Methodology: https://research.meta.ai/static/muse-spark-1-2-methodology |
| Terminal-Bench 2.1 | 76.2% | Default | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | 1 Oct 2026 | mini-SWE-agent | TB 2.1 (89 tasks, archived); effort not listed |
| Terminal-Bench 2.1 | 80.0% | Defaulteffort not stated | Maker's own figureVendor-reportedMeta ↗ | 9 Jul 2026 | — | Harness not stated; Meta's 1.2 post shows 76.2% with mini-swe-agent. Effort not stated on page (Meta's later comparisons run 1.1 at xhigh). Evaluation report link currently broken. |
| Terminal-Bench 4.0 | 6.1% | Extra highMuse Spark 1.1 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug muse-spark-1-1) (AA marks this variant deprecated) |
| Toolathlon-Verified | 75.6% | Defaulteffort not stated | Maker's own figureVendor-reportedMeta ↗ | 9 Jul 2026 | — | Effort not stated on page (Meta's later comparisons run 1.1 at xhigh). Evaluation report link currently broken. |
| Vals CorpFin v2 | 71.3% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 12 Aug 2026 | — | vals id meta/muse_spark_1_1; rank 4/134; ±0.889 stderr; $0.10923/test |
| Vals Finance Agent v2 | 57.2% | Defaulteffort not stated | Maker's own figureVendor-reportedMeta ↗ | 9 Jul 2026 | — | Effort not stated on page (Meta's later comparisons run 1.1 at xhigh). Evaluation report link currently broken. |
| Vals TaxEval v2 | 79.7% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | vals id meta/muse_spark_1_1; rank 2/145; ±0.77 stderr; $0.014921/test |