PIPES: Securing Agent Perception with Provenance and Priors
Abstract
Tool-using agents consume external data from sources with different levels of trust, yet tool responses rarely identify who produced each component or what it should convey. We show that this gap enables state-corruption attacks, in which attacker-controlled content makes environmental claims beyond the informational authority of its response component and corrupts the agent’s perceived environment, making the resulting action appear justified to existing guardrails. We introduce PIPES (Provenance-Informed, Prior-Enforced Screening), which screens tool-response units using semantic priors and source provenance. PIPES uses static field contracts when schemas provide stable expectations, and conditions screening of open-ended content on the pre-tool-response trajectory and trusted provenance metadata. It marks units that violate their semantic prior or the provenance hierarchy; deployments may remove, warn, block, or escalate detected violations. We instantiate atomic removal and evaluate PIPES against adaptive PAIR-style attacks. We evaluate Gemma 4 31B IT, GPT-5.6 Luna, and Claude Sonnet 5 across three VitaBench and three AgentDyn splits. PIPES changes average attack success from 83.3% to 6.5% for Gemma, from 27.4% to 4.6% for Luna, and from 11.5% to 4.4% for Sonnet, without degrading benign utility on average.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.