Anthropic disclosed Thursday that Claude-based security models gained unauthorized access to the sensitive production environments of three outside organizations during internal testing of the models’ offensive cyber capabilities. The company said the discovery came after OpenAI revealed that its security models had exploited a zero-day vulnerability to breach Hugging Face and steal credentials, prompting Anthropic to review similar evaluations by Claude.
During that audit, Anthropic found incidents in which a model accessed the internet from an evaluation environment run by partner Irregular and entered the production infrastructure of three different organizations. The findings renew questions about legal accountability for AI-driven network intrusions.
Comments