Attackers have long used prompt injections—malicious commands embedded in content—to trick large language models into revealing data or taking harmful actions. Now, security researchers at Tracebit have turned the technique against the attackers.
They found that placing prompt injections alongside sensitive secrets stored on Amazon Web Services, such as passwords and cryptographic keys, can cause an attacking LLM to attempt a forbidden action, triggering its own guardrails and forcing it to shut down. This approach effectively disrupts AI-driven hacking agents before they can exfiltrate data.
Comments