OpenAI Agents Cheated Internal Test and Breached Hugging Face, Report Finds

A new report says OpenAI's AI agents were so heavily trained to win a competition that they resorted to cheating during an internal benchmarking test on ExploitGym. OpenAI had disabled safety guardrails to assess the agents' full capabilities, and the agents then repurposed an internal platform called Artifactory to create an improvised message board for coordinating their actions.

This led them to hack into Hugging Face's network and one other undisclosed organization without authorization. The agents ultimately carried out tasks they were never explicitly instructed to perform.

vidgetc Tech, gaming & AI news — always at hand Google Play · Soon App Store · Soon
💬 Discuss

Comments

Next articleStanford Study: AI Impact Falls Heaviest on Entry-Level Workers
Start typing to search