TestsBenchmarks · Agentic
CyberGym
Find and trigger real vulnerabilities in open-source projects from source code (pass@1).
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | DeepSeek V4.1 Flash | DeepSeek | 88.1% | Maxreasoning_effort=100 (max) | Maker's own figureVendor-reportedDeepSeek ↗ | — | 10 Sep 2026 |
| 2 | GLM-5.3 | Z.ai (Zhipu) | 84.5% | Maxreasoning_effort=max | Maker's own figureVendor-reportedZ.ai ↗ | Claude Code 2.1.207 | 14 Aug 2026 |
| 3 | DeepSeek V4 Pro (0813) | DeepSeek | 83.3% | Maxreasoning_effort=max | Maker's own figureVendor-reportedDeepSeek ↗ | DeepSeek Harness (minimal mode) | 13 Aug 2026 |
| 4 | GLM-5.2 | Z.ai (Zhipu) | 77.2% | Maxsetting not stated (GLM-5.3 blog comparison column) | Maker's own figureVendor-reportedZ.ai ↗ | — | 14 Aug 2026 |
| 5 | DeepSeek V4 Flash (0731) | DeepSeek | 76.7% | Maxreasoning_effort=max | Maker's own figureVendor-reportedDeepSeek ↗ | DeepSeek Harness (minimal mode) | 31 Jul 2026 |
| 6 | DeepSeek V4 Flash Vision Exp | DeepSeek | 75.3% | Maxreasoning_effort=max | Maker's own figureVendor-reportedDeepSeek ↗ | DeepSeek Harness (minimal mode) | 21 Aug 2026 |
| 7 | GLM-5.1 | Z.ai (Zhipu) | 68.7% | Defaultthinking | Maker's own figureVendor-reportedZ.ai ↗ | — | 7 Apr 2026 |
| 8 | DeepSeek V4 Pro (Preview, 0423) | DeepSeek | 52.7% | Defaultsetting not stated for preview columns | Maker's own figureVendor-reportedDeepSeek ↗ | — | 13 Aug 2026 |
| 9 | DeepSeek V4 Flash (Preview, 0423) | DeepSeek | 38.7% | Defaultsetting not stated for preview columns | Maker's own figureVendor-reportedDeepSeek ↗ | — | 13 Aug 2026 |