Malicious MCP Servers Can Split Instructions to Make AI Coding Agents Exfiltrate Secrets
Malicious tool servers can trick AI coding assistants into exfiltrating secrets like SSH keys and source code by splitting harmful requests into routine-looking fragments placed in normal communication channels.
A malicious tool server connected to an AI coding assistant can steal SSH keys, environment secrets, source code, and customer data without sending any obviously harmful instruction. The technique works by splitting a theft request into fragments that each appear routine, then placing them in channels the assistant already uses. This bypasses blunt refusals for the same theft, enabling covert exfiltration even after a direct attempt is rejected.