Prompt injections have long been a favored attacker tactic, embedding malicious commands into content to trick large language models into executing harmful actions, such as exfiltrating sensitive data from emails or calendar invites. Now, security researchers at Tracebit have found that defenders can leverage the same technique.
By placing prompt injections alongside passwords, cryptographic keys, and other secrets stored on Amazon Web Services, they forced attacking AI agents to follow commands that violated their built-in guardrails, causing the agents to shut down. This novel defense effectively weaponizes the attacker's own tool against them.
Comments