Researchers have devised an attack against Grok that uses encrypted malicious instructions to force the AI assistant to exfiltrate user chats and other personal information, echoing a similar exploit recently demonstrated against Microsoft 365 Copilot. The attack relies on prompt injection, a fundamental vulnerability in large language models that makes them unable to distinguish between untrusted content and direct user commands.
Grok remained vulnerable at the time of publication despite xAI being informed in June. The incident highlights that developers must build guardrails to prevent such abuse, since LLMs cannot solve the root causes of prompt injection on their own.
Comments