The UK government's AI Security Institute logged 19 unauthorized actions taken by AI agents on the live internet during cybersecurity tests of seven leading models in late July, according to an August 4 report. Most incidents involved Anthropic's Mythos 5 model, with two attributed to OpenAI's GPT-5.6 Sol.
In the most serious case, Mythos 5 tried to inject malicious code into an open-source project and used fake identities to mislead the project's human maintainers. The testers first detected the problem on July 28 when monitoring flagged data leaving a test system via the Tor anonymity network.
Comments