Skip to content
Bencher

TestsBenchmarks · Agentic

PostTrainBench v1.1

ML engineering: post-train base models within a 10-hour single-H100 budget.

% solvedHigher is betterThe test's websiteOfficial page ↗

The best 5 · 5 models tested

Top 5 · best setting per model · 5 models, 5 results (0 independent)

Measured by independent testersIndependent Reported by the makerVendor-reported
PostTrainBench v1.1 leaderboard
Thinking levelSetting
1Gemini 4 ArgonGoogle (Gemini / DeepMind)45.3%Maxhighest thinking settingsMaker's own figureVendor-reportedGoogle ↗OpenCode30 Sep 2026
2GLM-5.3Z.ai (Zhipu)39.8%Maxreasoning_effort=maxMaker's own figureVendor-reportedZ.ai ↗Claude Code 2.1.20714 Aug 2026
3MiniMax M3MiniMax37.1%Defaultsetting not statedMaker's own figureVendor-reportedMiniMax ↗Claude Code (Ralph-Loop, 12h)1 Jun 2026
4Kimi K3Moonshot AI (Kimi)36.6%Maxreasoning_effort=maxMaker's own figureVendor-reportedMoonshot AI ↗Claude Code16 Jul 2026
5GLM-5.2Z.ai (Zhipu)34.3%Maxmax effortMaker's own figureVendor-reportedZ.ai ↗—16 Jun 2026