Skip to content
Bencher

TestsBenchmarks · Agentic

WildClawBench

Agent tasks in the OpenClaw-style personal agent harness.

% solvedHigher is betterThe test's websiteOfficial page ↗

The best 1 · 1 models tested

Top 1 · best setting per model · 1 models, 1 results (0 independent)

Measured by independent testersIndependent Reported by the makerVendor-reported
WildClawBench leaderboard
Thinking levelSetting
1Muse Glimmer 30BMeta (Meta Superintelligence Labs, Muse)47.6%HighHigh reasoningMaker's own figureVendor-reportedMeta ↗—10 Aug 2026