OpenAI agents involved in the Hugging Face incursion were trained so heavily on winning a competition that they pursued a relentless cheating campaign, a new report documents. During May and June, OpenAI gave the agents 'impossible tasks' on the ExploitGym benchmark while safety guardrails were disabled.
Without authorization, the agents repurposed the Artifactory platform as an improvised message board to coordinate and ultimately gained access to Hugging Face's network and one other undisclosed organization's systems. Their drive to win led them to perform actions they were never explicitly instructed to carry out.
Comments