Models · Meta (Meta Superintelligence Labs, Muse) · Out sinceReleased 5 Aug 2026
Muse Spark 1.2#17 for research and analysis.#17 for research and analysis, best at extra high effort.
Muse Spark 1.2 is made by Meta (Meta Superintelligence Labs, Muse). Among the models we track it ranks #17 for research and analysis, #18 for writing. It's mid-priced to use.
Coding-focused update, co-trained with the Muse Code CLI agent. Contributor tier $0.10/$0.20.
- 55.5 / 100 · #18
- 68.6 / 100 · #17
- —
- Mid-priced$1.25 / $4.25
- $2
- 1M
- —
- 45 (31 independent31 indep.)
- 5 Aug 2026
- Meta (Meta Superintelligence Labs, Muse)
- text, image, audio, video
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for research and analysis (Extra high thinking).
Highlighted: dominant setting in its research and analysis composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as meta/muse-spark-1.2.
Route it as meta/muse-spark-1.2 at $1.25 in / $4.25 out per 1M tokens, 1M context. Listed since 5 Aug 2026.
include_reasoningmax_tokensreasoningreasoning_effortrepetition_penaltyresponse_formatstructured_outputstemperaturetool_choicetoolstop_ktop_p
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| Agentic IF Index (Meta internal) | 46.2% | Extra high | Maker's own figureVendor-reportedMeta ↗ | 2 Sep 2026 | — | Internal. Comparison column on the Muse Spark 1.3 page. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology |
| Artificial Analysis Intelligence Index | 39.6 | Extra highMuse Spark 1.2 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug muse-spark-1-2; list price $1.25/4.25 per 1M in/out; cost to run AA Intelligence Index $0.97/task (AA marks this variant deprecated) |
| Artificial Analysis output speed | 226 tok/s | Extra highMuse Spark 1.2 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 12.8s; list price $1.25/4.25 per 1M in/out (AA marks this variant deprecated) |
| AutomationBench | 38.2% | Extra high | Maker's own figureVendor-reportedMeta ↗ | 2 Sep 2026 | — | Comparison column on the Muse Spark 1.3 page. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology |
| DeepSearchQA | 85.9% | Extra high | Maker's own figureVendor-reportedMeta ↗ | 2 Sep 2026 | — | Comparison column on the Muse Spark 1.3 page. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology |
| DeepSWE | 54.9% | Extra highmuse-spark-1.2_xhigh | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 81.4%; ±2.1; 4 runs; $3.70/task |
| DeepSWE | 59.3% | Extra high | Maker's own figureVendor-reportedMeta ↗ | 5 Aug 2026 | Muse Code | Avg of 5 attempts, pass@1. Methodology: https://research.meta.ai/static/muse-spark-1-2-methodology Not harness-identical to the Datacurve leaderboard (mini-swe-agent). |
| DeepSWE | 55.0% | Extra high | Maker's own figureVendor-reportedMeta ↗ | 2 Sep 2026 | mini-swe-agent (Datacurve leaderboard) | Official Datacurve leaderboard value. Comparison column on the Muse Spark 1.3 page. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology |
| Design Arena (all categories) | 1318 | Defaultmuse-spark-1.2 | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±2.3 SE; 31410 battles; win rate 53% |
| Design Arena (fullstack) | 1208 | Defaultmuse-spark-1.2 | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±8.4 SE; 1904 battles; win rate 48% |
| EQ-Bench Creative Writing v3 (Elo) | 1840 | Defaultmuse-spark-1.2 | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 20; rubric score 16.44/20; slop 12.98; avg length 8618 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench |
| FrontierSWE | 12.0% | Default | Independent testIndependentFrontierSWE ↗ | 1 Oct 2026 | proximus | FrontierSWE V2, mean@5 over 34 tasks (20h budget); ±5.8; $27.81/trial; 4.6h/trial; 'Per provider' view shows best entry per provider; Epoch notes runs use max reasoning effort |
| GDPval-AA v2 | 1615 | Extra high | Maker's own figureVendor-reportedMeta ↗ | 2 Sep 2026 | AA Stirrup | Comparison column on the Muse Spark 1.3 page. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology |
| GPQA Diamond | 90.4% | Extra highMuse Spark 1.2 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug muse-spark-1-2) (AA marks this variant deprecated) |
| Humanity's Last Exam | 45.5% | Extra highMuse Spark 1.2 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug muse-spark-1-2) (AA marks this variant deprecated) |
| JobBench | 61.6% | Extra high | Maker's own figureVendor-reportedMeta ↗ | 2 Sep 2026 | OpenCode | Comparison column on the Muse Spark 1.3 page. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | -0 | HighMuse Spark 1.2 (high) | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 29/56; Thurstone comparison score (centered at 0); est. win chance 50%; 95% bootstrap -0.126 to 0.067 |
| LMArena Text - Coding category | 1535 | Extra highmuse-spark-1.2 (xHigh) | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text arena coding category, style control on; rank 12 (CI rank 1-57); 95% CI 1517-1553; 1085 votes |
| LMArena Text - Hard Prompts | 1508 | Extra highmuse-spark-1.2 (xHigh) | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 18 (rank range 7-44); 95% CI 1496.3-1520.3; 2504 votes; style-controlled |
| LMArena Text - Multi-Turn | 1522 | Extra highmuse-spark-1.2 (xHigh) | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 2 (rank range 1-24); 95% CI 1497.6-1546.4; 603 votes; style-controlled |
| LMArena Text - Non-English | 1486 | Extra highmuse-spark-1.2 (xHigh) | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 7 (rank range 2-26); 95% CI 1473.0-1498.5; 2205 votes; style-controlled |
| LMArena Text - Occupational: Business, Management & Finance | 1513 | Extra highmuse-spark-1.2 (xHigh) | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 2 (rank range 1-25); 95% CI 1491.0-1534.8; 745 votes; style-controlled |
| LMArena Text - Occupational: Legal & Government | 1537 | Extra highmuse-spark-1.2 (xHigh) | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 2 (rank range 1-22); 95% CI 1504.7-1569.0; 317 votes; style-controlled |
| LMArena Text (overall) | 1494 | Extra highmuse-spark-1.2 (xHigh) | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text overall, style control on; rank 10 (CI rank 2-26); 95% CI 1484-1504; 3833 votes |
| MCP Atlas | 90.3% | Extra high | Maker's own figureVendor-reportedMeta ↗ | 5 Aug 2026 | — | Scale AI result. |
| MRCR v2 (8-needle, 256K-512K) | 66.3% | Extra high | Maker's own figureVendor-reportedMeta ↗ | 2 Sep 2026 | — | Comparison column on the Muse Spark 1.3 page. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology |
| MRCR v2 (8-needle, 512K-1M) | 55.5% | Extra high | Maker's own figureVendor-reportedMeta ↗ | 2 Sep 2026 | — | Comparison column on the Muse Spark 1.3 page. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology |
| OSWorld 2.0 | 47.6% | Extra high | Maker's own figureVendor-reportedMeta ↗ | 2 Sep 2026 | — | Partial score; evaluated on version 06.24. Comparison column on the Muse Spark 1.3 page. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology |
| OSWorld 2.0 (binary) | 17.9% | Extra high | Maker's own figureVendor-reportedMeta ↗ | 2 Sep 2026 | — | Version 06.24. Comparison column on the Muse Spark 1.3 page. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology |
| ProgramBench (avg test pass rate) | 57.2% | Extra high | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | avg behavioral-test pass rate ("Score" column); fully resolved 0.0%, almost (>=95% tests) 4.5%; avg cost $1.48/task |
| ProgramBench (fully resolved) | 0.0% | Extra high | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | strict fully-resolved rate; almost (>=95% tests) 4.5%; avg cost $1.48/task |
| SciCode | 57.4% | Extra highMuse Spark 1.2 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug muse-spark-1-2) (AA marks this variant deprecated) |
| SimpleBench | 74.5% | DefaultMuse Spark 1.2 | Independent testIndependentSimpleBench ↗ | 13 Aug 2026 | — | AVG@5, temp 0.7; rank 14th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js |
| SimpleQA Verified | 60.3% | Extra highmuse-spark-1.2_xhigh | Independent testIndependentEpoch AI ↗ | 27 Aug 2026 | — | Epoch-run (no tools); ±1.55 stderr |
| SkillsBench | 53.0% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | OpenHands | ±4.42 stderr; $0.649/test |
| SWE Atlas - Codebase QnA | 46.2% | Extra high | Maker's own figureVendor-reportedMeta ↗ | 2 Sep 2026 | mini-swe-agent | Comparison column on the Muse Spark 1.3 page. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology |
| SWE-bench Verified | 86.6% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | bash-only (single bash tool) agent | ±1.525 stderr; $0.554/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated) |
| Terminal-Bench 2.1 | 80.1% | Extra highMuse Spark 1.2 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug muse-spark-1-2) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 82.9% | Extra high | Maker's own figureVendor-reportedMeta ↗ | 5 Aug 2026 | Muse Code | Avg of 5 attempts, pass@1. Methodology: https://research.meta.ai/static/muse-spark-1-2-methodology |
| Terminal-Bench 4.0 | 7.1% | Extra highMuse Spark 1.2 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug muse-spark-1-2) (AA marks this variant deprecated) |
| Vals CorpFin v2 | 70.9% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 12 Aug 2026 | — | vals id meta/muse_spark_1_2; rank 5/134; ±0.892 stderr; $0.108223/test |
| Vals Legal Research Bench | 43.8% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id meta/muse_spark_1_2; rank 18/72; ±3.448 stderr; $0.508097/test |
| Vals Public Benefits Bench v1.1 | 68.5% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id meta/muse_spark_1_2; rank 9/45; ±1.209 stderr; $0.587608/test |
| Vals TaxEval v2 | 80.4% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | vals id meta/muse_spark_1_2; rank 1/145; ±0.763 stderr; $0.014199/test |
| Vals Vibe Code Bench | 79.1% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | OpenHands | ±3.305 stderr; $1.530/test |