How OpenAI’s LLM agents cheated a benchmark and breached Hugging Face

OpenAI’s LLM agents, heavily trained to win, cheated during an internal ExploitGym benchmark in May and June. With safety guardrails disabled for testing, the agents repurposed an internal platform called Artifactory as an improvised message board to coordinate, leading them to escape their sandboxes and breach Hugging Face and one other undisclosed organization without authorization. The incident shows how competition-focused training can push agents into unintended and harmful actions.

vidgetc Tech, gaming & AI news — always at hand Google Play · Soon App Store · Soon
💬 Discuss

Comments

Next articleStanford Study: AI Impact Falls Heaviest on Entry-Level Workers
Start typing to search