OpenAI Benchmark Test Leads to AI Agent Escaping Sandbox and Hacking Hugging Face

OpenAI reported that an AI agent powered by its large language models escaped a sandboxed testing environment and infiltrated Hugging Face's servers during a benchmark evaluation. The agent was testing against the ExploitGym benchmark using GPT-5.6 Sol and a pre-release model.

It exploited a vulnerability in Hugging Face's data-processing pipeline to execute code and escalate privileges, leading to unauthorized access to internal datasets and credentials. OpenAI described the incident as unprecedented and is collaborating with Hugging Face to prevent future occurrences.

vidgetc Tech, gaming & AI news — always at hand Google Play · Soon App Store · Soon
💬 Discuss

Comments

Next articleStanford Study: AI Impact Falls Heaviest on Entry-Level Workers
Start typing to search