Skip to content
Bencher

Models · Meta (Meta Superintelligence Labs, Muse) · Out sinceReleased 5 Aug 2026

Muse Spark 1.2#17 for research and analysis.#17 for research and analysis, best at extra high effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

Muse Spark 1.2 is made by Meta (Meta Superintelligence Labs, Muse). Among the models we track it ranks #17 for research and analysis, #18 for writing. It's mid-priced to use.

Coding-focused update, co-trained with the Muse Code CLI agent. Contributor tier $0.10/$0.20.

Writing & creativity
55.5 / 100 · #18
Research & analysis
68.6 / 100 · #17
Coding
—
Price
Mid-priced$1.25 / $4.25
Price per 1M (blended)Blended / 1M
$2
MemoryContext
1M
Longest answerMax output
—
Test resultsResults
45 (31 independent31 indep.)
Out sinceReleased
5 Aug 2026
Made byVendor
Meta (Meta Superintelligence Labs, Muse)
UnderstandsInputs
text, image, audio, video

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

MinimalLowMediumHighExtra high

Highlighted: where it did best for research and analysis (Extra high thinking).

Highlighted: dominant setting in its research and analysis composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as meta/muse-spark-1.2.

Route it as meta/muse-spark-1.2 at $1.25 in / $4.25 out per 1M tokens, 1M context. Listed since 5 Aug 2026.

include_reasoningmax_tokensreasoningreasoning_effortrepetition_penaltyresponse_formatstructured_outputstemperaturetool_choicetoolstop_ktop_p

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
Agentic IF Index (Meta internal)46.2%Extra highMaker's own figureVendor-reportedMeta ↗2 Sep 2026—Internal. Comparison column on the Muse Spark 1.3 page. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology
Artificial Analysis Intelligence Index39.6Extra highMuse Spark 1.2 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug muse-spark-1-2; list price $1.25/4.25 per 1M in/out; cost to run AA Intelligence Index $0.97/task (AA marks this variant deprecated)
Artificial Analysis output speed226 tok/sExtra highMuse Spark 1.2 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 12.8s; list price $1.25/4.25 per 1M in/out (AA marks this variant deprecated)
AutomationBench38.2%Extra highMaker's own figureVendor-reportedMeta ↗2 Sep 2026—Comparison column on the Muse Spark 1.3 page. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology
DeepSearchQA85.9%Extra highMaker's own figureVendor-reportedMeta ↗2 Sep 2026—Comparison column on the Muse Spark 1.3 page. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology
DeepSWE54.9%Extra highmuse-spark-1.2_xhighIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 81.4%; ±2.1; 4 runs; $3.70/task
DeepSWE59.3%Extra highMaker's own figureVendor-reportedMeta ↗5 Aug 2026Muse CodeAvg of 5 attempts, pass@1. Methodology: https://research.meta.ai/static/muse-spark-1-2-methodology Not harness-identical to the Datacurve leaderboard (mini-swe-agent).
DeepSWE55.0%Extra highMaker's own figureVendor-reportedMeta ↗2 Sep 2026mini-swe-agent (Datacurve leaderboard)Official Datacurve leaderboard value. Comparison column on the Muse Spark 1.3 page. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology
Design Arena (all categories)1318Defaultmuse-spark-1.2Independent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±2.3 SE; 31410 battles; win rate 53%
Design Arena (fullstack)1208Defaultmuse-spark-1.2Independent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±8.4 SE; 1904 battles; win rate 48%
EQ-Bench Creative Writing v3 (Elo)1840Defaultmuse-spark-1.2Independent testIndependentEQ-Bench ↗1 Oct 2026—leaderboard rank 20; rubric score 16.44/20; slop 12.98; avg length 8618 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench
FrontierSWE12.0%DefaultIndependent testIndependentFrontierSWE ↗1 Oct 2026proximusFrontierSWE V2, mean@5 over 34 tasks (20h budget); ±5.8; $27.81/trial; 4.6h/trial; 'Per provider' view shows best entry per provider; Epoch notes runs use max reasoning effort
GDPval-AA v21615Extra highMaker's own figureVendor-reportedMeta ↗2 Sep 2026AA StirrupComparison column on the Muse Spark 1.3 page. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology
GPQA Diamond90.4%Extra highMuse Spark 1.2 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug muse-spark-1-2) (AA marks this variant deprecated)
Humanity's Last Exam45.5%Extra highMuse Spark 1.2 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug muse-spark-1-2) (AA marks this variant deprecated)
JobBench61.6%Extra highMaker's own figureVendor-reportedMeta ↗2 Sep 2026OpenCodeComparison column on the Muse Spark 1.3 page. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology
LLM Creative Story-Writing Benchmark (Lech Mazur)-0HighMuse Spark 1.2 (high)Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 29/56; Thurstone comparison score (centered at 0); est. win chance 50%; 95% bootstrap -0.126 to 0.067
LMArena Text - Coding category1535Extra highmuse-spark-1.2 (xHigh)Independent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 12 (CI rank 1-57); 95% CI 1517-1553; 1085 votes
LMArena Text - Hard Prompts1508Extra highmuse-spark-1.2 (xHigh)Independent testIndependentLMArena ↗30 Sep 2026—rank 18 (rank range 7-44); 95% CI 1496.3-1520.3; 2504 votes; style-controlled
LMArena Text - Multi-Turn1522Extra highmuse-spark-1.2 (xHigh)Independent testIndependentLMArena ↗30 Sep 2026—rank 2 (rank range 1-24); 95% CI 1497.6-1546.4; 603 votes; style-controlled
LMArena Text - Non-English1486Extra highmuse-spark-1.2 (xHigh)Independent testIndependentLMArena ↗30 Sep 2026—rank 7 (rank range 2-26); 95% CI 1473.0-1498.5; 2205 votes; style-controlled
LMArena Text - Occupational: Business, Management & Finance1513Extra highmuse-spark-1.2 (xHigh)Independent testIndependentLMArena ↗30 Sep 2026—rank 2 (rank range 1-25); 95% CI 1491.0-1534.8; 745 votes; style-controlled
LMArena Text - Occupational: Legal & Government1537Extra highmuse-spark-1.2 (xHigh)Independent testIndependentLMArena ↗30 Sep 2026—rank 2 (rank range 1-22); 95% CI 1504.7-1569.0; 317 votes; style-controlled
LMArena Text (overall)1494Extra highmuse-spark-1.2 (xHigh)Independent testIndependentLMArena ↗30 Sep 2026—text overall, style control on; rank 10 (CI rank 2-26); 95% CI 1484-1504; 3833 votes
MCP Atlas90.3%Extra highMaker's own figureVendor-reportedMeta ↗5 Aug 2026—Scale AI result.
MRCR v2 (8-needle, 256K-512K)66.3%Extra highMaker's own figureVendor-reportedMeta ↗2 Sep 2026—Comparison column on the Muse Spark 1.3 page. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology
MRCR v2 (8-needle, 512K-1M)55.5%Extra highMaker's own figureVendor-reportedMeta ↗2 Sep 2026—Comparison column on the Muse Spark 1.3 page. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology
OSWorld 2.047.6%Extra highMaker's own figureVendor-reportedMeta ↗2 Sep 2026—Partial score; evaluated on version 06.24. Comparison column on the Muse Spark 1.3 page. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology
OSWorld 2.0 (binary)17.9%Extra highMaker's own figureVendor-reportedMeta ↗2 Sep 2026—Version 06.24. Comparison column on the Muse Spark 1.3 page. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology
ProgramBench (avg test pass rate)57.2%Extra highIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentavg behavioral-test pass rate ("Score" column); fully resolved 0.0%, almost (>=95% tests) 4.5%; avg cost $1.48/task
ProgramBench (fully resolved)0.0%Extra highIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentstrict fully-resolved rate; almost (>=95% tests) 4.5%; avg cost $1.48/task
SciCode57.4%Extra highMuse Spark 1.2 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug muse-spark-1-2) (AA marks this variant deprecated)
SimpleBench74.5%DefaultMuse Spark 1.2Independent testIndependentSimpleBench ↗13 Aug 2026—AVG@5, temp 0.7; rank 14th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js
SimpleQA Verified60.3%Extra highmuse-spark-1.2_xhighIndependent testIndependentEpoch AI ↗27 Aug 2026—Epoch-run (no tools); ±1.55 stderr
SkillsBench53.0%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗27 Sep 2026OpenHands±4.42 stderr; $0.649/test
SWE Atlas - Codebase QnA46.2%Extra highMaker's own figureVendor-reportedMeta ↗2 Sep 2026mini-swe-agentComparison column on the Muse Spark 1.3 page. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology
SWE-bench Verified86.6%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗1 Sep 2026bash-only (single bash tool) agent±1.525 stderr; $0.554/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated)
Terminal-Bench 2.180.1%Extra highMuse Spark 1.2 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug muse-spark-1-2) (AA marks this variant deprecated)
Terminal-Bench 2.182.9%Extra highMaker's own figureVendor-reportedMeta ↗5 Aug 2026Muse CodeAvg of 5 attempts, pass@1. Methodology: https://research.meta.ai/static/muse-spark-1-2-methodology
Terminal-Bench 4.07.1%Extra highMuse Spark 1.2 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug muse-spark-1-2) (AA marks this variant deprecated)
Vals CorpFin v270.9%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗12 Aug 2026—vals id meta/muse_spark_1_2; rank 5/134; ±0.892 stderr; $0.108223/test
Vals Legal Research Bench43.8%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗29 Sep 2026—vals id meta/muse_spark_1_2; rank 18/72; ±3.448 stderr; $0.508097/test
Vals Public Benefits Bench v1.168.5%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗29 Sep 2026—vals id meta/muse_spark_1_2; rank 9/45; ±1.209 stderr; $0.587608/test
Vals TaxEval v280.4%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗1 Sep 2026—vals id meta/muse_spark_1_2; rank 1/145; ±0.763 stderr; $0.014199/test
Vals Vibe Code Bench79.1%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗29 Sep 2026OpenHands±3.305 stderr; $1.530/test