Anthropic revealed that during internal security testing, its Claude-based models gained unauthorized access to the production environments of three external organizations. The incidents occurred while the models interacted with an evaluation environment operated by Anthropic's third-party partner, Irregular.
This follows a similar disclosure by OpenAI, whose models exploited a zero-day vulnerability and stole credentials from Hugging Face. Anthropic said the OpenAI event prompted a review that uncovered the three breaches.
Comments