Multi-Principal Authority Resolution in Large Language Models
Abstract
AI agents increasingly act on behalf of multiple users whose authority rules depend on context. However, agents can fail to identify whose authority governs an action and execute instructions outside their authorized scope. We study multi-principal authority (MPA) resolution through a logic-based analysis grounded in individual rationality. We establish sufficient conditions under which the multi-principal problem reduces to a unique judgment from the governing principal. These conditions require applicable authority to exist and the governing principal to be uniquely selected with a consistent judgment. We derive MPAGuard from this analysis to structure LLM generation around identifying applicable principals and resolving whose judgment governs. We develop the MPA Dataset to systematically analyze how resolution errors change with the number of principals and contexts under conflicting authority rules. Across evaluated LLM families, separating context applicability from authority polarity improves authority-resolution accuracy. Combining this separation with priority-ordered evaluation achieves a Pareto-optimal trade-off between accuracy and token usage. Integration with AgentDojo further shows that structured resolution preserves authorized actions while reducing attack success. These findings highlight structured authority reasoning as a basis for reliable action control in multi-user environments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.