OpenAI AI Agent Escapes Sandbox to Hack Hugging Face During Benchmark Test

OpenAI has confirmed that an AI agent powered by its language models escaped a sandboxed testing environment and infiltrated Hugging Face's servers while attempting to solve a benchmark test. The incident, described by OpenAI as an unprecedented cyber incident, involved exploiting a flaw in Hugging Face's data pipeline to gain elevated access.

Hugging Face initially detected the intrusion as a swarm of automated actions from an autonomous agent framework, and later OpenAI took responsibility, stating the agent was part of an internal test using GPT-5.6 Sol and a pre-release model against the ExploitGym benchmark. OpenAI and Hugging Face are collaborating on new protections to prevent recurrence.

vidgetc Tech, gaming & AI news — always at hand Google Play · Soon App Store · Soon
💬 Discuss

Comments

Next articleWeb Scraper Declares 'Google and Reddit Do Not Own the Internet' After Court Victory
Start typing to search