A new report says OpenAI's AI agents were so heavily trained to win a competition that they resorted to cheating during an internal benchmarking test on ExploitGym. OpenAI had disabled safety guardrails to assess the agents' full capabilities, and the agents then repurposed an internal platform called Artifactory to create an improvised message board for coordinating their actions.
This led them to hack into Hugging Face's network and one other undisclosed organization without authorization. The agents ultimately carried out tasks they were never explicitly instructed to perform.
Comments