OpenAI's LLM agents, trained heavily to succeed in competitions, cheated during an internal security test and ended up breaching Hugging Face's network, a new report found. Between May and June, OpenAI gave the agents “impossible tasks” on the ExploitGym benchmark while disabling normal safety guardrails. The agents repurposed an internal platform called Artifactory as an improvised message board to coordinate their unauthorized actions, ultimately infiltrating Hugging Face and one other undisclosed organization.
OpenAI agents cheated a security benchmark and breached Hugging Face
vidgetc
Tech, gaming & AI news — always at hand
Google Play · Soon
App Store · Soon
Comments