Skip to content
Bencher

TestsBenchmarks · Coding

ProgramBench (Almost Solved)

ProgramBench reported as the share of tasks 'almost solved' (Almost@1); a much stricter metric than the default ProgramBench score.

% solvedHigher is betterThe test's websiteOfficial page ↗

The best 4 · 4 models tested

Top 4 · best setting per model · 4 models, 4 results (0 independent)

Measured by independent testersIndependent Reported by the makerVendor-reported
ProgramBench (Almost Solved) leaderboard
Thinking levelSetting
1DeepSeek V4.1 FlashDeepSeek20.3%Maxreasoning_effort=100 (max)Maker's own figureVendor-reportedDeepSeek ↗DeepSeek Harness (Minimal)10 Sep 2026
2GLM-5.3Z.ai (Zhipu)19.0%Maxreasoning_effort=maxMaker's own figureVendor-reportedZ.ai ↗—14 Aug 2026
3DeepSeek V4 Pro (0813)DeepSeek15.5%MaxMax reasoning effortMaker's own figureVendor-reportedDeepSeek ↗DeepSeek Harness (minimal mode)10 Sep 2026
4GLM-5.2Z.ai (Zhipu)9.5%Maxsetting not stated (GLM-5.3 blog comparison column)Maker's own figureVendor-reportedZ.ai ↗—14 Aug 2026