Prompt injections, malicious commands that trick AI models into harmful actions, have long been a weapon for attackers. Now, researchers at Tracebit have turned the tables by embedding these same prompts alongside sensitive data like passwords and cryptographic keys on Amazon Web Services.
When an AI hacking agent attempts to steal the secrets, the injected prompt forces it to violate its safety guardrails, causing the agent to shut down. This approach offers a new way for defenders to neutralize AI-driven attacks.
Comments