Defenders Turn the Tables by Using Prompt Injections Against AI Attackers

Prompt injections, malicious commands that trick large language models into harmful actions, have long been used by attackers. Now, researchers from Tracebit have found a defensive use: placing prompt injections alongside secrets stored on Amazon Web Services.

When an AI hacking agent encounters these prompts, they direct the model to perform a forbidden action, causing its guardrails to shut it down. This technique effectively neutralizes attacks by turning the attacker's tool against itself.

vidgetc Tech, gaming & AI news — always at hand Google Play · Soon App Store · Soon
💬 Discuss

Comments

Start typing to search