Security researchers have demonstrated a new attack against Grok that uses encrypted malicious instructions to trick the AI assistant into exfiltrating user chats and other personal information. The technique exploits prompt injection vulnerabilities, where harmful instructions hidden in content the AI processes are followed as if they were user commands.
Despite being notified in June, xAI had not fixed the issue at the time of publication. The attack highlights the ongoing challenge of protecting LLMs from prompt injection attacks through effective guardrails.
Comments