Attackers have long used prompt injections—malicious commands embedded in content—to trick large language models into leaking data or performing harmful actions. Now, researchers from Tracebit have found that defenders can use the same technique by placing prompt injections alongside sensitive secrets stored on Amazon Web Services.
When an AI hacking agent attempts to retrieve those secrets, the injected command directs the LLM to violate its own safety guardrails, causing it to shut down. This approach effectively weaponizes prompt injections as a defense mechanism against automated AI attacks.
Comments