Models · Meta (Meta Superintelligence Labs, Muse) · Out sinceReleased 8 Apr 2026
Muse Spark
Muse Spark is made by Meta (Meta Superintelligence Labs, Muse). We don't have enough test results yet to rank it.
Original Muse Spark (pre-1.1); not on OpenRouter.
- —
- —
- —
- Price unknown— / —
- —
- —
- —
- 28 (16 independent16 indep.)
- 8 Apr 2026
- Meta (Meta Superintelligence Labs, Muse)
- text, image
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
You can't choose how long this one thinks.
No adjustable reasoning setting listed.
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| BabyVision | 39.9% | Defaulteffort not stated | Maker's own figureVendor-reportedMeta ↗ | 9 Jul 2026 | — | 'Muse Spark' comparison column on the Muse Spark 1.1 page. |
| CharXiv Reasoning (no tools) | 88.9% | Defaulteffort not stated | Maker's own figureVendor-reportedMeta ↗ | 9 Jul 2026 | — | 'Muse Spark' comparison column on the Muse Spark 1.1 page. |
| DeepSWE | 10.0% | Defaulteffort not stated | Maker's own figureVendor-reportedMeta ↗ | 9 Jul 2026 | — | 'Muse Spark' comparison column on the Muse Spark 1.1 page. |
| FrontierScience Research | 38.0% | MaxContemplating mode | Maker's own figureVendor-reportedMeta ↗ | 8 Apr 2026 | — | Integer as published. |
| Humanity's Last Exam | 58.0% | MaxContemplating mode | Maker's own figureVendor-reportedMeta ↗ | 8 Apr 2026 | — | Contemplating mode = multiple agents reasoning in parallel; post does not say whether tools were used. Integer as published. |
| Humanity's Last Exam | 40.6% | DefaultMuse Spark | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 4 (Scale rank accounts for CI); ±1.92; entry added 2026-04-08 |
| Humanity's Last Exam (with tools) | 50.4% | Defaulteffort not stated | Maker's own figureVendor-reportedMeta ↗ | 9 Jul 2026 | — | With tools. 'Muse Spark' comparison column on the Muse Spark 1.1 page. |
| JobBench | 17.0% | Defaulteffort not stated | Maker's own figureVendor-reportedMeta ↗ | 9 Jul 2026 | — | 'Muse Spark' comparison column on the Muse Spark 1.1 page. |
| LMArena Text - Coding category | 1530 | Defaultmuse-spark | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text arena coding category, style control on; rank 22 (CI rank 7-48); 95% CI 1520-1540; 3930 votes |
| LMArena Text - Hard Prompts | 1506 | Defaultmuse-spark | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 20 (rank range 9-41); 95% CI 1499.3-1513.1; 9145 votes; style-controlled |
| LMArena Text - Multi-Turn | 1494 | Defaultmuse-spark | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 18 (rank range 6-55); 95% CI 1480.4-1506.6; 2284 votes; style-controlled |
| LMArena Text - Non-English | 1477 | Defaultmuse-spark | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 18 (rank range 6-33); 95% CI 1470.0-1484.8; 7451 votes; style-controlled |
| LMArena Text - Occupational: Business, Management & Finance | 1489 | Defaultmuse-spark | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 13 (rank range 3-41); 95% CI 1477.7-1500.8; 2866 votes; style-controlled |
| LMArena Text - Occupational: Legal & Government | 1501 | Defaultmuse-spark | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 12 (rank range 1-60); 95% CI 1481.8-1520.4; 1016 votes; style-controlled |
| LMArena Text (overall) | 1489 | Defaultmuse-spark | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text overall, style control on; rank 15 (CI rank 5-27); 95% CI 1483-1495; 14128 votes |
| MCP Atlas | 82.2% | DefaultMuse Spark | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 2 (Scale rank accounts for CI); ±2.3; entry added 2026-04-08 |
| MCP Atlas | 82.2% | Defaulteffort not stated | Maker's own figureVendor-reportedMeta ↗ | 9 Jul 2026 | — | 'Muse Spark' comparison column on the Muse Spark 1.1 page. |
| OSWorld-Verified | 53.3% | Defaulteffort not stated | Maker's own figureVendor-reportedMeta ↗ | 9 Jul 2026 | — | 'Muse Spark' comparison column on the Muse Spark 1.1 page. |
| PRBench Finance (Scale) | 52.4% | DefaultMuse Spark | Independent testIndependentScale AI (SEAL) ↗ | 8 Apr 2026 | — | Scale rank 5; ±0.06 CI |
| PRBench Legal (Scale) | 52.3% | DefaultMuse Spark | Independent testIndependentScale AI (SEAL) ↗ | 8 Apr 2026 | — | Scale rank 3; ±0.06 CI |
| SWE Atlas - Codebase QnA | 24.2% | DefaultMuse Spark | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 13 (Scale rank accounts for CI); ±4.6; entry added 2026-04-08 |
| SWE Atlas - Test Writing | 31.1% | DefaultMuse Spark | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 5 (Scale rank accounts for CI); ±5.76; entry added 2026-04-08 |
| SWE-Bench Pro (private/commercial set) | 44.7% | DefaultMuse Spark* | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 3 (Scale rank accounts for CI); ±6.05; entry added 2026-04-08; * = Scale footnote (see page) |
| SWE-Bench Pro (public, v1) | 55.0% | DefaultMuse Spark* | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 3 (Scale rank accounts for CI); ±3.6; entry added 2026-04-08; * = Scale footnote (see page) |
| SWE-Bench Pro (public, v1) | 55.0% | Defaulteffort not stated | Maker's own figureVendor-reportedMeta ↗ | 9 Jul 2026 | — | 'Muse Spark' comparison column on the Muse Spark 1.1 page. |
| Terminal-Bench 2.1 | 67.3% | Defaulteffort not stated | Maker's own figureVendor-reportedMeta ↗ | 9 Jul 2026 | — | 'Muse Spark' comparison column on the Muse Spark 1.1 page. |
| Toolathlon-Verified | 49.4% | Defaulteffort not stated | Maker's own figureVendor-reportedMeta ↗ | 9 Jul 2026 | — | 'Muse Spark' comparison column on the Muse Spark 1.1 page. |
| Vals TaxEval v2 | 77.7% | Defaultvals id meta/muse_spark | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | vals id meta/muse_spark; rank 3/145; ±0.806 stderr; $0.001922/test |