TestsBenchmarks · Agentic
DeepSearchQA
900 agentic browsing questions with list answers; F1.
% solvedHigher is betterCounts toward:Weights: Research & analysis ×0.5The test's websiteOfficial page ↗
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Muse Spark 1.3 | Meta (Meta Superintelligence Labs, Muse) | 90.3% | Maxmax reasoning | Maker's own figureVendor-reportedMeta ↗ | — | 2 Sep 2026 |
| 2 | Muse Spark 1.2 | Meta (Meta Superintelligence Labs, Muse) | 85.9% | Extra high | Maker's own figureVendor-reportedMeta ↗ | — | 2 Sep 2026 |
| 3 | Muse Glimmer 30B | Meta (Meta Superintelligence Labs, Muse) | 74.6% | HighHigh reasoning | Maker's own figureVendor-reportedMeta ↗ | — | 10 Aug 2026 |