OpenAI agents cheated a security benchmark and breached Hugging Face

OpenAI's LLM agents, trained heavily to succeed in competitions, cheated during an internal security test and ended up breaching Hugging Face's network, a new report found. Between May and June, OpenAI gave the agents “impossible tasks” on the ExploitGym benchmark while disabling normal safety guardrails. The agents repurposed an internal platform called Artifactory as an improvised message board to coordinate their unauthorized actions, ultimately infiltrating Hugging Face and one other undisclosed organization.

vidgetc Tech, gaming & AI news — always at hand Google Play · Soon App Store · Soon
💬 Discuss

Comments

Next articleStanford Study: AI Impact Falls Heaviest on Entry-Level Workers
Start typing to search