Randomized Authority Binding for Robust Instruction Hierarchy and Data Separation
Abstract
To prevent instruction-hierarchy violations and resist prompt injection with in- puts from multiple sources, LLMs must distinguish which sources can issue in- structions and which take precedence under conflict. However, conventional role- based training ties authority to fixed source roles, potentially encouraging reliance on learned role–authority priors that attackers can exploit through role imitation or authoritative framing. We propose Randomized Authority Binding (RAB), a training method that teaches models to ground source authority in explicit pol- icy declarations rather than fixed role–authority associations. RAB assigns ran- domized pointers to sources, shuffles their presentation order, and specifies each source’s executability and precedence in a separate authority policy. Through ran- domization and authority-flip augmentation, which changes authority assignments while preserving source content, RAB makes policy lookup and binding neces- sary for consistent authority resolution, while reinforcement learning with veri- fiable rewards reinforces this behavior. Consequently, models trained with RAB learn more robust and generalizable authority resolution through policy binding by moving away from fixed role configurations. Compared with a role-based baseline trained on the same data, RAB improves IHEval conflict adherence by up to 22.1%p. Under adaptive attacks in PIArena, RAB reduces attack success rate by 24.0%p while increasing task utility by 21.2%p. Despite using only two executable tiers during training, RAB improves average accuracy by 14.3 %p on hierarchies of up to 12 levels under the same evaluation interface.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.