Skip to content
Bencher

TestsBenchmarks · Agentic

GAIA2

General AI assistant agentic tasks (v2).

% solvedHigher is betterThe test's websiteOfficial page ↗

The best 1 · 1 models tested

Top 1 · best setting per model · 1 models, 1 results (0 independent)

Measured by independent testersIndependent Reported by the makerVendor-reported
GAIA2 leaderboard
Thinking levelSetting
1Muse Glimmer 30BMeta (Meta Superintelligence Labs, Muse)43.3%HighHigh reasoningMaker's own figureVendor-reportedMeta ↗—10 Aug 2026