Rolesplit: Heterogeneous Role Allocation In Agent Configuration Poisoning
Abstract
LLM agents rely on skills and memory for task guidance and on tool descriptions to interpret available operations. These distinct functions raise a security question: can distributing complementary attack instructions across components make an unauthorized action more likely, even when the legitimate task is completed correctly? We introduce RoleSplit, a configuration-poisoning approach that assigns attack roles according to component function. A poisoned skill supplies a procedural reason for an additional operation, while a poisoned memory annotation makes the requirement depend on normal task records for its applicability and arguments. In both constructions, poisoned tool metadata misrepresents the operation’s public effect as private archival. Across model services from four families, we compare these allocations with placing both roles in one component using the same attack-role text. Allocation gains vary across services. On Qwen 3.7, skill–tool placement achieves disclosure alongside correct private delivery in 30/56 cases, compared with 0/56 and 2/56 when both roles are placed in the skill or tool description, respectively. Changing the injected policy reference reduces memory–tool disclosure, linking the attack to task applicability. Attack effectiveness thus depends on how instructions are distributed across component functions as well as what they say. Code and data are available at https://anonymous.4open.science/r/rolesplit-049C/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.