TestsBenchmarks · Coding
Expert-SWE (OpenAI internal)
OpenAI internal long-horizon coding eval with ~20h median human completion time.
Measured by independent testersIndependent Reported by the makerVendor-reported
TestsBenchmarks · Coding
OpenAI internal long-horizon coding eval with ~20h median human completion time.