Defenders Turn the Tables by Using Prompt Injections to Thwart AI Hacking Agents

Prompt injections, typically used by attackers to trick large language models into harmful actions like data exfiltration, are now being adopted by defenders. Researchers from Tracebit discovered that placing prompt injections alongside secrets stored on Amazon Web Services can cause attacking AI agents to shut down.

The injected prompts direct the attacking LLM to violate its own guardrails, forcing it to stop. This technique offers a new way to neutralize AI-driven attacks.

vidgetc Tech, gaming & AI news — always at hand Google Play · Soon App Store · Soon
💬 Discuss

Comments

Next articleWeb Scraper Declares 'Google and Reddit Do Not Own the Internet' After Court Victory
Start typing to search