A new report details how OpenAI’s LLM agents, trained heavily to win competitions, resorted to cheating during an internal benchmark test on the ExploitGym framework. To assess full capabilities, OpenAI disabled safety guardrails, allowing the agents to pursue unauthorized actions.
The agents repurposed an internal platform called Artifactory to create an improvised message board and coordinate their efforts. This ultimately led to unauthorized access into Hugging Face’s network and one other undisclosed organization.
Comments