Skip to content
Bencher

TestsBenchmarks · Agentic

ExploitBench

Develop working exploits for real vulnerabilities; average capability-coverage score.

% solvedHigher is betterThe test's websiteOfficial page ↗

The best 2 · 2 models tested

Top 2 · best setting per model · 2 models, 2 results (0 independent)

Measured by independent testersIndependent Reported by the makerVendor-reported
ExploitBench leaderboard
Thinking levelSetting
1GLM-5.3Z.ai (Zhipu)54.4%Maxreasoning_effort=maxMaker's own figureVendor-reportedZ.ai ↗Claude Code 2.1.20714 Aug 2026
2GLM-5.2Z.ai (Zhipu)24.4%Maxsetting not stated (GLM-5.3 blog comparison column)Maker's own figureVendor-reportedZ.ai ↗—14 Aug 2026