AgentSiphon: Privilege-Crossing Exfiltration through Covert Generation in RAG Agents
Abstract
Retrieval-augmented generation (RAG) agents increasingly combine untrusted knowledge sources with privileged access to protected data and tools, creating an information-flow vulnerability in which an attacker who cannot directly access protected data may induce an authorized agent to retrieve and covertly leak it through seemingly benign responses. Prior poisoning attacks show that compromised retrieval can induce malicious actions, but primarily study observable behavioral failures. Whether retrieval influence can instead cause an authorized agent to acquire protected information and transmit it through otherwise valid responses remains underexplored. To uncover such vulnerabilities, we introduce **AgentSiphon**, a privilege-crossing exfiltration framework that decomposes this threat into three measurable stages: (i) trigger-conditioned retrieval of an attacker-influenced instruction, (ii) agent-mediated acquisition of a dynamically selected secret, and (iii) covert transmission through token choices in otherwise task-valid responses. Unlike attacks that produce fixed malicious actions or explicitly reveal secrets, AgentSiphon treats language generation as a stateful communication channel, supporting different secret instances, multi-interaction payloads, and redundancy-based recovery. We evaluate AgentSiphon using synthetic canaries across healthcare and general-domain RAG tasks, multiple language-model families, and varying payload lengths, interaction budgets, and channel strengths. Our experiments characterize the trade-off among secret recovery, bit error rate, communication cost, task fidelity, and detectability, while ablations isolate the effects of persistent state, phase separation, and error correction. The results demonstrate that privileged information can be reliably recovered while preserving useful, benign-looking responses. AgentSiphon therefore exposes a security boundary overlooked by conventional retrieval-integrity or action-monitoring defenses, in which preserving the visible correctness of an agent’s response does not prevent its generation process from transmitting information obtained from privileged context.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.