Prompt injections, typically used by attackers to trick large language models into harmful actions like data exfiltration, are now being adopted by defenders. Researchers from Tracebit discovered that placing prompt injections alongside secrets stored on Amazon Web Services can cause attacking AI agents to shut down.
The injected prompts direct the attacking LLM to violate its own guardrails, forcing it to stop. This technique offers a new way to neutralize AI-driven attacks.
Comments