Prompt injections, traditionally used by attackers to trick large language models into harmful actions, are now being adopted by defenders. Researchers from Tracebit demonstrated that embedding prompt injections alongside sensitive data like passwords and cryptographic keys on Amazon Web Services can effectively neutralize AI hacking agents. The injected prompts direct the attacking LLM to violate its own safety guardrails, causing it to shut down rather than execute malicious commands.
Defenders Turn Prompt Injections Against AI Hackers
vidgetc
Tech, gaming & AI news — always at hand
Google Play · Soon
App Store · Soon
Comments