Skip to content
Bencher

TestsBenchmarks · Agentic

CyberGym

Find and trigger real vulnerabilities in open-source projects from source code (pass@1).

% solvedHigher is betterThe test's websiteOfficial page ↗

The best 9 · 9 models tested

Top 9 · best setting per model · 9 models, 9 results (0 independent)

Measured by independent testersIndependent Reported by the makerVendor-reported
CyberGym leaderboard
Thinking levelSetting
1DeepSeek V4.1 FlashDeepSeek88.1%Maxreasoning_effort=100 (max)Maker's own figureVendor-reportedDeepSeek ↗—10 Sep 2026
2GLM-5.3Z.ai (Zhipu)84.5%Maxreasoning_effort=maxMaker's own figureVendor-reportedZ.ai ↗Claude Code 2.1.20714 Aug 2026
3DeepSeek V4 Pro (0813)DeepSeek83.3%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗DeepSeek Harness (minimal mode)13 Aug 2026
4GLM-5.2Z.ai (Zhipu)77.2%Maxsetting not stated (GLM-5.3 blog comparison column)Maker's own figureVendor-reportedZ.ai ↗—14 Aug 2026
5DeepSeek V4 Flash (0731)DeepSeek76.7%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗DeepSeek Harness (minimal mode)31 Jul 2026
6DeepSeek V4 Flash Vision ExpDeepSeek75.3%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗DeepSeek Harness (minimal mode)21 Aug 2026
7GLM-5.1Z.ai (Zhipu)68.7%DefaultthinkingMaker's own figureVendor-reportedZ.ai ↗—7 Apr 2026
8DeepSeek V4 Pro (Preview, 0423)DeepSeek52.7%Defaultsetting not stated for preview columnsMaker's own figureVendor-reportedDeepSeek ↗—13 Aug 2026
9DeepSeek V4 Flash (Preview, 0423)DeepSeek38.7%Defaultsetting not stated for preview columnsMaker's own figureVendor-reportedDeepSeek ↗—13 Aug 2026