Models · Meta (Meta Superintelligence Labs, Muse) · Out sinceReleased 2 Sep 2026
Muse Spark 1.3#3 for research and analysis.#3 for research and analysis, best at max effort.
Muse Spark 1.3 is made by Meta (Meta Superintelligence Labs, Muse). Among the models we track it ranks #3 for research and analysis, #12 for coding, #19 for writing. It's mid-priced to use.
Meta's current flagship (Meta Superintelligence Labs; Muse replaced Llama). 'max' effort is Standard tier only. Contributor tier muse-spark-1.3-contributor at $0.10/$0.20 (data used for training). Audio understanding 'not fully supported'. Max output 131072 from Meta's coding-agent config docs. OpenRouter vendor slug is 'meta' (not 'meta-llama'). Meta says Muse Spark open weights are on the roadmap.
- 54.6 / 100 · #19
- 84.0 / 100 · #3
- 55.7 / 100 · #12
- Mid-priced$1.25 / $4.25
- $2
- 1M
- 131K
- 93 (81 independent81 indep.)
- 2 Sep 2026
- Meta (Meta Superintelligence Labs, Muse)
- text, image, audio, video
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for research and analysis (Maximum thinking).
Highlighted: dominant setting in its research and analysis composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as meta/muse-spark-1.3.
Route it as meta/muse-spark-1.3 at $1.25 in / $4.25 out per 1M tokens, 1M context. Listed since 2 Sep 2026.
include_reasoningmax_tokensreasoningreasoning_effortrepetition_penaltyresponse_formatstructured_outputstemperaturetool_choicetoolstop_ktop_p
Thinking level
Reasoning effort
How long should Muse Spark 1.3Where Muse Spark 1.3think?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
Artificial Analysis Coding Agent Index
Best at MaxBest at Max
Terminal-Bench 4.0
Best at MaxBest at Max
CursorBench
Best at MaxBest at Max
DeepSWE
Best at Extra high, worse abovePeaks at Extra high
LMArena Code Arena (WebDev)
Best at MaxBest at Max
ProgramBench (avg test pass rate)
Best at MaxBest at Max
SWE Atlas - Codebase QnA
Best at MaxBest at Max
Vals Vibe Code Bench
Best at MaxBest at Max
ProgramBench (fully resolved)
Best at MaxBest at Max
Terminal-Bench 2.1
Best at Extra high, worse abovePeaks at Extra high
SciCode
Best at Extra high, worse abovePeaks at Extra high
Artificial Analysis Intelligence Index
Best at MaxBest at Max
Vals Index
Best at MaxBest at Max
AA-LCR
Best at Extra highBest at Extra high
AA-Omniscience Accuracy
Best at MaxBest at Max
AA-Omniscience Hallucination Rate
Best at Extra high, worse abovePeaks at Extra high
Lower is better on this test
AA-Omniscience Index
Best at MaxBest at Max
FrontierMath (Tiers 1-3)
Best at Extra high, worse abovePeaks at Extra high
GPQA Diamond
Best at Extra high, worse abovePeaks at Extra high
Humanity's Last Exam
Best at MaxBest at Max
Vals Legal Research Bench
Best at MaxBest at Max
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA-Briefcase v1.1 | 1586 | MaxMuse Spark 1.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-LCR | 83.0% | Extra highMuse Spark 1.3 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 83.0% | MaxMuse Spark 1.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-Omniscience Accuracy | 41.5% | Extra highMuse Spark 1.3 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 43.6% | MaxMuse Spark 1.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Hallucination Rate | 31.5% | Extra highMuse Spark 1.3 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 32.9% | MaxMuse Spark 1.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Index | 23.1 | Extra highMuse Spark 1.3 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 25 | MaxMuse Spark 1.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| Agentic IF Index (Meta internal) | 57.8% | Maxmax reasoning | Maker's own figureVendor-reportedMeta ↗ | 2 Sep 2026 | — | Internal. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology |
| Artificial Analysis Coding Agent Index | 48.3 | Extra highMuse Spark 1.3 (xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Muse Code | agent Muse Code 1.0.2 RC; components: DeepSWE v1.1 73.2, SWE-Atlas-QnA 54.6, Terminal-Bench v4 17.2; avg cost $3.47/task; avg wall time 18 min/task |
| Artificial Analysis Coding Agent Index | 54.3 | MaxMuse Spark 1.3 (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Muse Code | agent Muse Code; components: DeepSWE v1.1 71.7, SWE-Atlas-QnA 59.4, Terminal-Bench v4 31.8; avg cost $3.98/task; avg wall time 18 min/task |
| Artificial Analysis Intelligence Index | 45.1 | Extra highMuse Spark 1.3 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug muse-spark-1-3-xhigh; list price $1.25/4.25 per 1M in/out; cost to run AA Intelligence Index $1.37/task |
| Artificial Analysis Intelligence Index | 48.1 | MaxMuse Spark 1.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug muse-spark-1-3; list price $1.25/4.25 per 1M in/out; cost to run AA Intelligence Index $1.60/task |
| Artificial Analysis output speed | 139 tok/s | Extra highMuse Spark 1.3 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 32.6s; list price $1.25/4.25 per 1M in/out |
| Artificial Analysis output speed | 174 tok/s | MaxMuse Spark 1.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 36.0s; list price $1.25/4.25 per 1M in/out |
| AutomationBench | 49.6% | Maxmax reasoning | Maker's own figureVendor-reportedMeta ↗ | 2 Sep 2026 | — | Public v3 task set, pass@1. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology |
| CursorBench | 24.3% | MinimalMuse Spark 1.3 Minimal | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 60; $0.56/task; 10,620 tokens/task; 34 steps/task |
| CursorBench | 29.3% | LowMuse Spark 1.3 Low | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 53; $0.93/task; 17,483 tokens/task; 47 steps/task |
| CursorBench | 32.6% | MediumMuse Spark 1.3 Medium | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 46; $1.49/task; 27,255 tokens/task; 64 steps/task |
| CursorBench | 33.4% | HighMuse Spark 1.3 High | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 42; $1.66/task; 30,654 tokens/task; 69 steps/task |
| CursorBench | 37.5% | Extra highMuse Spark 1.3 Extra High | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 32; $2.10/task; 40,891 tokens/task; 83 steps/task |
| CursorBench | 41.6% | MaxMuse Spark 1.3 Max | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 23; $2.64/task; 52,005 tokens/task; 98 steps/task |
| DeepSearchQA | 90.3% | Maxmax reasoning | Maker's own figureVendor-reportedMeta ↗ | 2 Sep 2026 | — | F1. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology |
| DeepSWE | 73.2% | Extra highMuse Spark 1.3 (xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Muse Code | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| DeepSWE | 71.7% | MaxMuse Spark 1.3 (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Muse Code | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| DeepSWE | 75.4% | Maxmax reasoning | Maker's own figureVendor-reportedMeta ↗ | 2 Sep 2026 | mini-swe-agent | Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology |
| Design Arena (all categories) | 1358 | Maxmuse-spark-1.3-max | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±2.7 SE; 21275 battles; win rate 57.7% |
| Design Arena (all categories) | 1359 | Defaultmuse-spark-1.3 | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±2.5 SE; 25856 battles; win rate 58.2% |
| Design Arena (fullstack) | 1310 | Maxmuse-spark-1.3-max | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±5.5 SE; 5657 battles; win rate 60.7% |
| Design Arena (fullstack) | 1293 | Defaultmuse-spark-1.3 | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±6.3 SE; 4006 battles; win rate 56.2% |
| Epoch Capabilities Index | 156.9 | Default | Independent testIndependentEpoch AI ↗ | 1 Oct 2026 | — | Epoch Capabilities Index; 90% CI 154.7-159.6; best of listed model versions |
| EQ-Bench Creative Writing v3 (Elo) | 1906 | Default*muse-spark-1.3 | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 16; rubric score 16.70/20; slop 10.73; avg length 8001 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench; marked new (*) |
| FrontierMath (Tiers 1-3) | 74.4% | Extra highmuse-spark-1.3_xhigh | Independent testIndependentEpoch AI ↗ | 16 Sep 2026 | — | FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.6pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv) |
| FrontierMath (Tiers 1-3) | 74.0% | Maxmuse-spark-1.3_max | Independent testIndependentEpoch AI ↗ | 18 Sep 2026 | — | FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.6pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv) |
| GDPval-AA v2 | 1754 | Maxmax reasoning | Maker's own figureVendor-reportedMeta ↗ | 2 Sep 2026 | AA Stirrup | Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology |
| GDPval-AA v2.1 | 1672 | MaxMuse Spark 1.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | 95% CI 1650.4-1694.25 |
| GPQA Diamond | 94.1% | Extra highMuse Spark 1.3 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug muse-spark-1-3-xhigh) |
| GPQA Diamond | 93.5% | MaxMuse Spark 1.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug muse-spark-1-3) |
| Harvey LAB-AA | 95.5% | Extra highMuse Spark 1.3 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | criteria pass rate, AA-run |
| HLE Diamond | 25.4% | Defaultmuse-spark-1.3 | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 8 (Scale rank accounts for CI); ±2.7; entry added 2026-03-10 |
| Humanity's Last Exam | 47.5% | Extra highMuse Spark 1.3 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug muse-spark-1-3-xhigh) |
| Humanity's Last Exam | 48.7% | MaxMuse Spark 1.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug muse-spark-1-3) |
| IOI (Vals) | 56.6% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±2.517 stderr; $1.730/test |
| JobBench | 64.9% | Maxmax reasoning | Maker's own figureVendor-reportedMeta ↗ | 2 Sep 2026 | OpenCode | Mean rubric score. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | 0.6 | HighMuse Spark 1.3 (high) | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 20/56; Thurstone comparison score (centered at 0); est. win chance 59%; 95% bootstrap 0.488 to 0.669 |
| LMArena Code Arena (WebDev) | 1623 | Extra highmuse-spark-1.3 (xHigh) | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | Code Arena | WebDev overall (agentic web-dev, raw); rank 18 (CI rank 14-24); 95% CI 1615-1632; 6078 votes |
| LMArena Code Arena (WebDev) | 1655 | Maxmuse-spark-1.3-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | Code Arena | WebDev overall (agentic web-dev, raw); rank 13 (CI rank 9-14); 95% CI 1647-1664; 7004 votes |
| LMArena Text - Coding category | 1539 | Maxmuse-spark-1.3-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text arena coding category, style control on; rank 10 (CI rank 1-31); 95% CI 1528-1550; 3142 votes |
| LMArena Text - Creative Writing | 1459 | Maxmuse-spark-1.3-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 27 (rank range 12-61); 95% CI 1445.9-1471.3; 2482 votes; style-controlled |
| LMArena Text - Expert | 1528 | Maxmuse-spark-1.3-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 13 (rank range 1-43); 95% CI 1511.7-1544.0; 1325 votes; style-controlled |
| LMArena Text - Hard Prompts | 1517 | Maxmuse-spark-1.3-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 10 (rank range 4-24); 95% CI 1509.6-1524.5; 7706 votes; style-controlled |
| LMArena Text - Instruction Following | 1486 | Maxmuse-spark-1.3-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 15 (rank range 6-34); 95% CI 1476.7-1495.5; 4263 votes; style-controlled |
| LMArena Text - Longer Query | 1500 | Maxmuse-spark-1.3-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 14 (rank range 7-33); 95% CI 1491.6-1509.1; 5391 votes; style-controlled |
| LMArena Text - Multi-Turn | 1491 | Maxmuse-spark-1.3-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 20 (rank range 7-64); 95% CI 1476.4-1505.0; 1775 votes; style-controlled |
| LMArena Text - Non-English | 1485 | Maxmuse-spark-1.3-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 8 (rank range 2-23); 95% CI 1477.6-1492.8; 7030 votes; style-controlled |
| LMArena Text - Occupational: Business, Management & Finance | 1505 | Maxmuse-spark-1.3-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 4 (rank range 1-23); 95% CI 1492.3-1518.0; 2184 votes; style-controlled |
| LMArena Text - Occupational: Legal & Government | 1506 | Maxmuse-spark-1.3-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 8 (rank range 1-49); 95% CI 1487.0-1524.5; 1003 votes; style-controlled |
| LMArena Text - Occupational: Writing, Literature & Language | 1471 | Maxmuse-spark-1.3-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 23 (rank range 8-46); 95% CI 1460.0-1481.9; 3187 votes; style-controlled |
| LMArena Text (overall) | 1495 | Maxmuse-spark-1.3-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text overall, style control on; rank 8 (CI rank 3-20); 95% CI 1488-1501; 11698 votes |
| MRCR v2 (8-needle, 256K-512K) | 98.5% | Maxmax reasoning | Maker's own figureVendor-reportedMeta ↗ | 2 Sep 2026 | — | Mean sequence-matcher ratio, 100 examples. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology |
| MRCR v2 (8-needle, 512K-1M) | 98.1% | Maxmax reasoning | Maker's own figureVendor-reportedMeta ↗ | 2 Sep 2026 | — | Mean sequence-matcher ratio, 100 examples. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology |
| OSWorld 2.0 | 66.9% | Maxmax reasoning | Maker's own figureVendor-reportedMeta ↗ | 2 Sep 2026 | — | Partial score; version 08.08. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology |
| OSWorld 2.0 (binary) | 32.0% | Maxmax reasoning | Maker's own figureVendor-reportedMeta ↗ | 2 Sep 2026 | — | Strict binary completion. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology |
| PRBench Finance (Scale) | 59.5% | DefaultMuse Spark 1.3 | Independent testIndependentScale AI (SEAL) ↗ | 14 Sep 2026 | — | Scale rank 1; ±1.66 CI |
| PRBench Legal (Scale) | 61.6% | DefaultMuse Spark 1.3 | Independent testIndependentScale AI (SEAL) ↗ | 14 Sep 2026 | — | Scale rank 1; ±1.72 CI |
| ProgramBench (avg test pass rate) | 68.6% | Extra high | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | avg behavioral-test pass rate ("Score" column); fully resolved 1.0%, almost (>=95% tests) 16.5%; avg cost $2.02/task |
| ProgramBench (avg test pass rate) | 70.8% | Max | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | avg behavioral-test pass rate ("Score" column); fully resolved 2.5%, almost (>=95% tests) 25.0%; avg cost $6.46/task |
| ProgramBench (fully resolved) | 1.0% | Extra high | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | strict fully-resolved rate; almost (>=95% tests) 16.5%; avg cost $2.02/task |
| ProgramBench (fully resolved) | 2.5% | Max | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | strict fully-resolved rate; almost (>=95% tests) 25.0%; avg cost $6.46/task |
| SciCode | 59.7% | Extra highMuse Spark 1.3 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug muse-spark-1-3-xhigh) |
| SciCode | 58.8% | MaxMuse Spark 1.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug muse-spark-1-3) |
| SimpleBench | 81.8% | DefaultMuse Spark 1.3 | Independent testIndependentSimpleBench ↗ | 3 Sep 2026 | — | AVG@5, temp 0.7; rank 7th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js |
| SWE Atlas - Codebase QnA | 54.6% | Extra highMuse Spark 1.3 (xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Muse Code | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| SWE Atlas - Codebase QnA | 59.4% | MaxMuse Spark 1.3 (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Muse Code | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| SWE Atlas - Codebase QnA | 59.4% | Maxmax reasoning | Maker's own figureVendor-reportedMeta ↗ | 2 Sep 2026 | mini-swe-agent | Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology |
| Terminal-Bench 2.1 | 85.4% | Extra highMuse Spark 1.3 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug muse-spark-1-3-xhigh) |
| Terminal-Bench 2.1 | 72.3% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | Terminus 2 | ±0.375 stderr; $0.663/test |
| Terminal-Bench 2.1 | 84.3% | MaxMuse Spark 1.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug muse-spark-1-3) |
| Terminal-Bench 2.1 | 79.0% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | Terminus 2 | ±0.991 stderr; $0.597/test |
| Terminal-Bench 2.1 | 88.8% | Maxmax reasoning | Maker's own figureVendor-reportedMeta ↗ | 2 Sep 2026 | native coding harness (Muse Code) | Mean pass@1. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology |
| Terminal-Bench 4.0 | 17.2% | Extra highMuse Spark 1.3 (xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Muse Code | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 16.7% | Extra highMuse Spark 1.3 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug muse-spark-1-3-xhigh) |
| Terminal-Bench 4.0 | 33.3% | MaxMuse Spark 1.3 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug muse-spark-1-3) |
| Terminal-Bench 4.0 | 31.8% | MaxMuse Spark 1.3 (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Muse Code | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 24.8% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±0.505 stderr; $6.655/test |
| Vals Code Migration | 47.4% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±4.265 stderr; $14.210/test |
| Vals Index | 53.2% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 30 Sep 2026 | — | ±1.066 stderr; $3.376/test |
| Vals Index | 58.2% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 30 Sep 2026 | — | ±1.188 stderr; $3.787/test |
| Vals Legal Research Bench | 40.9% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id meta/muse_spark_1_3; rank 23/72; ±3.417 stderr; $0.650063/test |
| Vals Legal Research Bench | 55.3% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id meta/muse_spark_1_3_max; rank 1/72; ±3.456 stderr; $0.588812/test |
| Vals Vibe Code Bench | 82.9% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | OpenHands | ±2.901 stderr; $2.102/test |
| Vals Vibe Code Bench | 85.9% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | OpenHands | ±2.508 stderr; $2.543/test |