Prompt injections, malicious commands that trick large language models into harmful actions, have long been used by attackers. Now, researchers from Tracebit have found a defensive use: placing prompt injections alongside secrets stored on Amazon Web Services.
When an AI hacking agent encounters these prompts, they direct the model to perform a forbidden action, causing its guardrails to shut it down. This technique effectively neutralizes attacks by turning the attacker's tool against itself.
Comments