Skip to content
Bencher

TestsBenchmarks · Agentic

DeepSearchQA

900 agentic browsing questions with list answers; F1.

% solvedHigher is betterCounts toward:Weights: Research & analysis ×0.5The test's websiteOfficial page ↗

The best 3 · 3 models tested

Top 3 · best setting per model · 3 models, 3 results (0 independent)

Measured by independent testersIndependent Reported by the makerVendor-reported
DeepSearchQA leaderboard
Thinking levelSetting
1Muse Spark 1.3Meta (Meta Superintelligence Labs, Muse)90.3%Maxmax reasoningMaker's own figureVendor-reportedMeta ↗—2 Sep 2026
2Muse Spark 1.2Meta (Meta Superintelligence Labs, Muse)85.9%Extra highMaker's own figureVendor-reportedMeta ↗—2 Sep 2026
3Muse Glimmer 30BMeta (Meta Superintelligence Labs, Muse)74.6%HighHigh reasoningMaker's own figureVendor-reportedMeta ↗—10 Aug 2026