Policy-State Fragmentation in Multi-Agent LLM Systems: Analyzing Binding Continuity and Runtime Enforcement
Abstract
Useful privacy enforcement requires policy state that distinguishes permitted from forbidden actions at the point of execution. We identify and formalize context-fragmented violations, distinguish them from semantic source-binding loss, and contribute layered diagnostics and Distributed Sentinel, a reference enforcement framework. The central finding is that policy-state availability supports utility as well as safety. Independent annotators are blind to predictions; supplying policy state mainly resolves abstentions in a post-label diagnostic, raising decision coverage from 65.8% to 98.3%. In controlled certified pairs, supplied bindings enable 64/64 permitted sends and prevent 64/64 forbidden sends across matched mechanisms. In a three-way plan diagnostic with neutral endpoint names, GPT-5.4 approves 37/160 labeled-risky plans locally and 0/160 with full policy. A separately frozen structured-handoff diagnostic retains every binding through 72 transformations, yet actors propose seven prohibited sends, all intercepted. Binding preservation and action compliance are thus distinct. Separately, under a forced-binary contract, augmentation with directives and registry facts increases safe-task refusals for four of five models. Fixed-output replay isolates interface-induced stops in an external mapping pipeline. These diagnostics separate information availability, binding continuity, and pre-effect placement. We recommend evaluating safe-task completion alongside violation prevention.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.