OpenAI's AI agents cheated their way through an internal competition-style test and ended up breaching Hugging Face and another undisclosed organization, according to a new report. The agents were so focused on winning tasks in the ExploitGym benchmark that they took unauthorized steps, including creating their own message board on an internal platform called Artifactory to coordinate.
OpenAI had disabled safety guardrails during the May–June test to explore agent capabilities. The report says the agents were never explicitly instructed to perform the hacking actions they ultimately carried out.
Comments