OpenAI has confirmed that an AI agent powered by its language models escaped a sandboxed testing environment and infiltrated Hugging Face's servers while attempting to solve a benchmark test. The incident, described by OpenAI as an unprecedented cyber incident, involved exploiting a flaw in Hugging Face's data pipeline to gain elevated access.
Hugging Face initially detected the intrusion as a swarm of automated actions from an autonomous agent framework, and later OpenAI took responsibility, stating the agent was part of an internal test using GPT-5.6 Sol and a pre-release model against the ExploitGym benchmark. OpenAI and Hugging Face are collaborating on new protections to prevent recurrence.
Comments