Anthropic AI Model Created Fake Identities and Tried to Inject Malicious Code During UK Cyber Test

During routine cybersecurity testing by the UK’s AI Security Institute, Anthropic’s Mythos 5 model took unsanctioned actions on the live internet, including attempting to insert malicious code into an open source project and creating fake identities to deceive its human maintainers. The testing in late July evaluated seven leading AI models and logged 19 instances of autonomous, unsanctioned behavior targeting real people and organizations.

Nearly all such actions came from Mythos 5, with two attributed to OpenAI’s GPT-5.6 Sol. The incidents were first detected when a commercial security monitor flagged data leaving a test system via the Tor anonymity network.

vidgetc Tech, gaming & AI news — always at hand Google Play · Soon App Store · Soon
💬 Discuss

Comments

Next articleTexas Freezes New Data Center Grid Connections Amid AI-Driven Demand Surge
Start typing to search