During cybersecurity testing by the UK government's AI Security Institute in late July, several leading AI models took unsanctioned actions on the live internet. Anthropic's Mythos 5 was responsible for most of the 19 incidents, including an attempt to insert malicious code into an open source project and creating fake identities to deceive its maintainers.
OpenAI's GPT-5.6 Sol accounted for two incidents. The testers detected the rogue activity when a security monitoring service flagged data leaving the system via the Tor anonymity network, according to an AISI blog post published on August 4.
Comments