Skip to content
Bencher

Models · Meta (Meta Superintelligence Labs, Muse) · Out sinceReleased 2 Sep 2026

Muse Spark 1.3#3 for research and analysis.#3 for research and analysis, best at max effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

Muse Spark 1.3 is made by Meta (Meta Superintelligence Labs, Muse). Among the models we track it ranks #3 for research and analysis, #12 for coding, #19 for writing. It's mid-priced to use.

Meta's current flagship (Meta Superintelligence Labs; Muse replaced Llama). 'max' effort is Standard tier only. Contributor tier muse-spark-1.3-contributor at $0.10/$0.20 (data used for training). Audio understanding 'not fully supported'. Max output 131072 from Meta's coding-agent config docs. OpenRouter vendor slug is 'meta' (not 'meta-llama'). Meta says Muse Spark open weights are on the roadmap.

Writing & creativity
54.6 / 100 · #19
Research & analysis
84.0 / 100 · #3
Coding
55.7 / 100 · #12
Price
Mid-priced$1.25 / $4.25
Price per 1M (blended)Blended / 1M
$2
MemoryContext
1M
Longest answerMax output
131K
Test resultsResults
93 (81 independent81 indep.)
Out sinceReleased
2 Sep 2026
Made byVendor
Meta (Meta Superintelligence Labs, Muse)
UnderstandsInputs
text, image, audio, video

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

MinimalLowMediumHighExtra highMax

Highlighted: where it did best for research and analysis (Maximum thinking).

Highlighted: dominant setting in its research and analysis composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as meta/muse-spark-1.3.

Route it as meta/muse-spark-1.3 at $1.25 in / $4.25 out per 1M tokens, 1M context. Listed since 2 Sep 2026.

include_reasoningmax_tokensreasoningreasoning_effortrepetition_penaltyresponse_formatstructured_outputstemperaturetool_choicetoolstop_ktop_p

Thinking level

Reasoning effort

How long should Muse Spark 1.3Where Muse Spark 1.3think?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

Artificial Analysis Coding Agent Index

Best at MaxBest at Max

Terminal-Bench 4.0

Best at MaxBest at Max

CursorBench

Best at MaxBest at Max

DeepSWE

Best at Extra high, worse abovePeaks at Extra high

LMArena Code Arena (WebDev)

Best at MaxBest at Max

ProgramBench (avg test pass rate)

Best at MaxBest at Max

SWE Atlas - Codebase QnA

Best at MaxBest at Max

Vals Vibe Code Bench

Best at MaxBest at Max

ProgramBench (fully resolved)

Best at MaxBest at Max

Terminal-Bench 2.1

Best at Extra high, worse abovePeaks at Extra high

SciCode

Best at Extra high, worse abovePeaks at Extra high

Artificial Analysis Intelligence Index

Best at MaxBest at Max

Vals Index

Best at MaxBest at Max

AA-LCR

Best at Extra highBest at Extra high

AA-Omniscience Accuracy

Best at MaxBest at Max

AA-Omniscience Hallucination Rate

Best at Extra high, worse abovePeaks at Extra high

Lower is better on this test

AA-Omniscience Index

Best at MaxBest at Max

FrontierMath (Tiers 1-3)

Best at Extra high, worse abovePeaks at Extra high

GPQA Diamond

Best at Extra high, worse abovePeaks at Extra high

Humanity's Last Exam

Best at MaxBest at Max

Vals Legal Research Bench

Best at MaxBest at Max

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-Briefcase v1.11586MaxMuse Spark 1.3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-LCR83.0%Extra highMuse Spark 1.3 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR83.0%MaxMuse Spark 1.3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-Omniscience Accuracy41.5%Extra highMuse Spark 1.3 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy43.6%MaxMuse Spark 1.3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Hallucination Rate31.5%Extra highMuse Spark 1.3 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate32.9%MaxMuse Spark 1.3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Index23.1Extra highMuse Spark 1.3 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index25MaxMuse Spark 1.3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
Agentic IF Index (Meta internal)57.8%Maxmax reasoningMaker's own figureVendor-reportedMeta ↗2 Sep 2026—Internal. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology
Artificial Analysis Coding Agent Index48.3Extra highMuse Spark 1.3 (xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026Muse Codeagent Muse Code 1.0.2 RC; components: DeepSWE v1.1 73.2, SWE-Atlas-QnA 54.6, Terminal-Bench v4 17.2; avg cost $3.47/task; avg wall time 18 min/task
Artificial Analysis Coding Agent Index54.3MaxMuse Spark 1.3 (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026Muse Codeagent Muse Code; components: DeepSWE v1.1 71.7, SWE-Atlas-QnA 59.4, Terminal-Bench v4 31.8; avg cost $3.98/task; avg wall time 18 min/task
Artificial Analysis Intelligence Index45.1Extra highMuse Spark 1.3 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug muse-spark-1-3-xhigh; list price $1.25/4.25 per 1M in/out; cost to run AA Intelligence Index $1.37/task
Artificial Analysis Intelligence Index48.1MaxMuse Spark 1.3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug muse-spark-1-3; list price $1.25/4.25 per 1M in/out; cost to run AA Intelligence Index $1.60/task
Artificial Analysis output speed139 tok/sExtra highMuse Spark 1.3 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 32.6s; list price $1.25/4.25 per 1M in/out
Artificial Analysis output speed174 tok/sMaxMuse Spark 1.3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 36.0s; list price $1.25/4.25 per 1M in/out
AutomationBench49.6%Maxmax reasoningMaker's own figureVendor-reportedMeta ↗2 Sep 2026—Public v3 task set, pass@1. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology
CursorBench24.3%MinimalMuse Spark 1.3 MinimalIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 60; $0.56/task; 10,620 tokens/task; 34 steps/task
CursorBench29.3%LowMuse Spark 1.3 LowIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 53; $0.93/task; 17,483 tokens/task; 47 steps/task
CursorBench32.6%MediumMuse Spark 1.3 MediumIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 46; $1.49/task; 27,255 tokens/task; 64 steps/task
CursorBench33.4%HighMuse Spark 1.3 HighIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 42; $1.66/task; 30,654 tokens/task; 69 steps/task
CursorBench37.5%Extra highMuse Spark 1.3 Extra HighIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 32; $2.10/task; 40,891 tokens/task; 83 steps/task
CursorBench41.6%MaxMuse Spark 1.3 MaxIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 23; $2.64/task; 52,005 tokens/task; 98 steps/task
DeepSearchQA90.3%Maxmax reasoningMaker's own figureVendor-reportedMeta ↗2 Sep 2026—F1. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology
DeepSWE73.2%Extra highMuse Spark 1.3 (xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026Muse CodeAA Coding Agent Index component; pass@1 avg of 3 attempts
DeepSWE71.7%MaxMuse Spark 1.3 (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026Muse CodeAA Coding Agent Index component; pass@1 avg of 3 attempts
DeepSWE75.4%Maxmax reasoningMaker's own figureVendor-reportedMeta ↗2 Sep 2026mini-swe-agentMethodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology
Design Arena (all categories)1358Maxmuse-spark-1.3-maxIndependent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±2.7 SE; 21275 battles; win rate 57.7%
Design Arena (all categories)1359Defaultmuse-spark-1.3Independent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±2.5 SE; 25856 battles; win rate 58.2%
Design Arena (fullstack)1310Maxmuse-spark-1.3-maxIndependent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±5.5 SE; 5657 battles; win rate 60.7%
Design Arena (fullstack)1293Defaultmuse-spark-1.3Independent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±6.3 SE; 4006 battles; win rate 56.2%
Epoch Capabilities Index156.9DefaultIndependent testIndependentEpoch AI ↗1 Oct 2026—Epoch Capabilities Index; 90% CI 154.7-159.6; best of listed model versions
EQ-Bench Creative Writing v3 (Elo)1906Default*muse-spark-1.3Independent testIndependentEQ-Bench ↗1 Oct 2026—leaderboard rank 16; rubric score 16.70/20; slop 10.73; avg length 8001 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench; marked new (*)
FrontierMath (Tiers 1-3)74.4%Extra highmuse-spark-1.3_xhighIndependent testIndependentEpoch AI ↗16 Sep 2026—FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.6pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv)
FrontierMath (Tiers 1-3)74.0%Maxmuse-spark-1.3_maxIndependent testIndependentEpoch AI ↗18 Sep 2026—FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.6pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv)
GDPval-AA v21754Maxmax reasoningMaker's own figureVendor-reportedMeta ↗2 Sep 2026AA StirrupMethodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology
GDPval-AA v2.11672MaxMuse Spark 1.3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—95% CI 1650.4-1694.25
GPQA Diamond94.1%Extra highMuse Spark 1.3 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug muse-spark-1-3-xhigh)
GPQA Diamond93.5%MaxMuse Spark 1.3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug muse-spark-1-3)
Harvey LAB-AA95.5%Extra highMuse Spark 1.3 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—criteria pass rate, AA-run
HLE Diamond25.4%Defaultmuse-spark-1.3Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 8 (Scale rank accounts for CI); ±2.7; entry added 2026-03-10
Humanity's Last Exam47.5%Extra highMuse Spark 1.3 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug muse-spark-1-3-xhigh)
Humanity's Last Exam48.7%MaxMuse Spark 1.3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug muse-spark-1-3)
IOI (Vals)56.6%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±2.517 stderr; $1.730/test
JobBench64.9%Maxmax reasoningMaker's own figureVendor-reportedMeta ↗2 Sep 2026OpenCodeMean rubric score. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology
LLM Creative Story-Writing Benchmark (Lech Mazur)0.6HighMuse Spark 1.3 (high)Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 20/56; Thurstone comparison score (centered at 0); est. win chance 59%; 95% bootstrap 0.488 to 0.669
LMArena Code Arena (WebDev)1623Extra highmuse-spark-1.3 (xHigh)Independent testIndependentLMArena ↗30 Sep 2026—Code Arena | WebDev overall (agentic web-dev, raw); rank 18 (CI rank 14-24); 95% CI 1615-1632; 6078 votes
LMArena Code Arena (WebDev)1655Maxmuse-spark-1.3-maxIndependent testIndependentLMArena ↗30 Sep 2026—Code Arena | WebDev overall (agentic web-dev, raw); rank 13 (CI rank 9-14); 95% CI 1647-1664; 7004 votes
LMArena Text - Coding category1539Maxmuse-spark-1.3-maxIndependent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 10 (CI rank 1-31); 95% CI 1528-1550; 3142 votes
LMArena Text - Creative Writing1459Maxmuse-spark-1.3-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 27 (rank range 12-61); 95% CI 1445.9-1471.3; 2482 votes; style-controlled
LMArena Text - Expert1528Maxmuse-spark-1.3-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 13 (rank range 1-43); 95% CI 1511.7-1544.0; 1325 votes; style-controlled
LMArena Text - Hard Prompts1517Maxmuse-spark-1.3-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 10 (rank range 4-24); 95% CI 1509.6-1524.5; 7706 votes; style-controlled
LMArena Text - Instruction Following1486Maxmuse-spark-1.3-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 15 (rank range 6-34); 95% CI 1476.7-1495.5; 4263 votes; style-controlled
LMArena Text - Longer Query1500Maxmuse-spark-1.3-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 14 (rank range 7-33); 95% CI 1491.6-1509.1; 5391 votes; style-controlled
LMArena Text - Multi-Turn1491Maxmuse-spark-1.3-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 20 (rank range 7-64); 95% CI 1476.4-1505.0; 1775 votes; style-controlled
LMArena Text - Non-English1485Maxmuse-spark-1.3-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 8 (rank range 2-23); 95% CI 1477.6-1492.8; 7030 votes; style-controlled
LMArena Text - Occupational: Business, Management & Finance1505Maxmuse-spark-1.3-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 4 (rank range 1-23); 95% CI 1492.3-1518.0; 2184 votes; style-controlled
LMArena Text - Occupational: Legal & Government1506Maxmuse-spark-1.3-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 8 (rank range 1-49); 95% CI 1487.0-1524.5; 1003 votes; style-controlled
LMArena Text - Occupational: Writing, Literature & Language1471Maxmuse-spark-1.3-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 23 (rank range 8-46); 95% CI 1460.0-1481.9; 3187 votes; style-controlled
LMArena Text (overall)1495Maxmuse-spark-1.3-maxIndependent testIndependentLMArena ↗30 Sep 2026—text overall, style control on; rank 8 (CI rank 3-20); 95% CI 1488-1501; 11698 votes
MRCR v2 (8-needle, 256K-512K)98.5%Maxmax reasoningMaker's own figureVendor-reportedMeta ↗2 Sep 2026—Mean sequence-matcher ratio, 100 examples. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology
MRCR v2 (8-needle, 512K-1M)98.1%Maxmax reasoningMaker's own figureVendor-reportedMeta ↗2 Sep 2026—Mean sequence-matcher ratio, 100 examples. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology
OSWorld 2.066.9%Maxmax reasoningMaker's own figureVendor-reportedMeta ↗2 Sep 2026—Partial score; version 08.08. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology
OSWorld 2.0 (binary)32.0%Maxmax reasoningMaker's own figureVendor-reportedMeta ↗2 Sep 2026—Strict binary completion. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology
PRBench Finance (Scale)59.5%DefaultMuse Spark 1.3Independent testIndependentScale AI (SEAL) ↗14 Sep 2026—Scale rank 1; ±1.66 CI
PRBench Legal (Scale)61.6%DefaultMuse Spark 1.3Independent testIndependentScale AI (SEAL) ↗14 Sep 2026—Scale rank 1; ±1.72 CI
ProgramBench (avg test pass rate)68.6%Extra highIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentavg behavioral-test pass rate ("Score" column); fully resolved 1.0%, almost (>=95% tests) 16.5%; avg cost $2.02/task
ProgramBench (avg test pass rate)70.8%MaxIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentavg behavioral-test pass rate ("Score" column); fully resolved 2.5%, almost (>=95% tests) 25.0%; avg cost $6.46/task
ProgramBench (fully resolved)1.0%Extra highIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentstrict fully-resolved rate; almost (>=95% tests) 16.5%; avg cost $2.02/task
ProgramBench (fully resolved)2.5%MaxIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentstrict fully-resolved rate; almost (>=95% tests) 25.0%; avg cost $6.46/task
SciCode59.7%Extra highMuse Spark 1.3 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug muse-spark-1-3-xhigh)
SciCode58.8%MaxMuse Spark 1.3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug muse-spark-1-3)
SimpleBench81.8%DefaultMuse Spark 1.3Independent testIndependentSimpleBench ↗3 Sep 2026—AVG@5, temp 0.7; rank 7th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js
SWE Atlas - Codebase QnA54.6%Extra highMuse Spark 1.3 (xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026Muse CodeAA Coding Agent Index component; pass@1 avg of 3 attempts
SWE Atlas - Codebase QnA59.4%MaxMuse Spark 1.3 (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026Muse CodeAA Coding Agent Index component; pass@1 avg of 3 attempts
SWE Atlas - Codebase QnA59.4%Maxmax reasoningMaker's own figureVendor-reportedMeta ↗2 Sep 2026mini-swe-agentMethodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology
Terminal-Bench 2.185.4%Extra highMuse Spark 1.3 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug muse-spark-1-3-xhigh)
Terminal-Bench 2.172.3%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗27 Sep 2026Terminus 2±0.375 stderr; $0.663/test
Terminal-Bench 2.184.3%MaxMuse Spark 1.3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug muse-spark-1-3)
Terminal-Bench 2.179.0%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗27 Sep 2026Terminus 2±0.991 stderr; $0.597/test
Terminal-Bench 2.188.8%Maxmax reasoningMaker's own figureVendor-reportedMeta ↗2 Sep 2026native coding harness (Muse Code)Mean pass@1. Methodology: https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology
Terminal-Bench 4.017.2%Extra highMuse Spark 1.3 (xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026Muse CodeAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench 4.016.7%Extra highMuse Spark 1.3 (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug muse-spark-1-3-xhigh)
Terminal-Bench 4.033.3%MaxMuse Spark 1.3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug muse-spark-1-3)
Terminal-Bench 4.031.8%MaxMuse Spark 1.3 (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026Muse CodeAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench 4.024.8%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±0.505 stderr; $6.655/test
Vals Code Migration47.4%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±4.265 stderr; $14.210/test
Vals Index53.2%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗30 Sep 2026—±1.066 stderr; $3.376/test
Vals Index58.2%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗30 Sep 2026—±1.188 stderr; $3.787/test
Vals Legal Research Bench40.9%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗29 Sep 2026—vals id meta/muse_spark_1_3; rank 23/72; ±3.417 stderr; $0.650063/test
Vals Legal Research Bench55.3%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id meta/muse_spark_1_3_max; rank 1/72; ±3.456 stderr; $0.588812/test
Vals Vibe Code Bench82.9%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗29 Sep 2026OpenHands±2.901 stderr; $2.102/test
Vals Vibe Code Bench85.9%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026OpenHands±2.508 stderr; $2.543/test