Provenance Is Not Influence: The Semantic-Information Flow Gap in LLM Agents
Abstract
Large Language Model (LLM) agents can transform untrusted external information through neural regeneration before using it to act, potentially decoupling semantic influence from the provenance observable at the authorization interface. We term this decoupling the Semantic-Information Flow Gap (SIFG): source-derived, action-relevant information can persist even when its source provenance is no longer observable. We characterize SIFG by separately measuring security-relevant semantic survival, authorization-visible provenance retention, and counterfactual action dependence, and show that this decoupling persists across the evaluated models, agent workflows, action settings, and transmission channels. We then formalize an observational boundary for provenance-only enforcement: when an attack and a benign execution induce the same gate-visible observation, no rule restricted to that observation can distinguish them. In the evaluated settings, provenance-only policies therefore face a security–utility tension, while additional provenance signals do not expose the missing distinction and an action-semantic signal can expose it. These results suggest that secure authorization after neural regeneration should treat provenance as distinct from semantic influence and incorporate policy-relevant signals beyond provenance alone. Our code is available at https://anonymous.4open.science/r/SIFG.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.