acceptodds
Under review as a conference paper at ICLR 2027

Multi-Principal Authority Resolution in Large Language Models

Abstract

AI agents increasingly act on behalf of multiple users whose authority rules depend on context. However, agents can fail to identify whose authority governs an action and execute instructions outside their authorized scope. We study multi-principal authority (MPA) resolution through a logic-based analysis grounded in individual rationality. We establish sufficient conditions under which the multi-principal problem reduces to a unique judgment from the governing principal. These conditions require applicable authority to exist and the governing principal to be uniquely selected with a consistent judgment. We derive MPAGuard from this analysis to structure LLM generation around identifying applicable principals and resolving whose judgment governs. We develop the MPA Dataset to systematically analyze how resolution errors change with the number of principals and contexts under conflicting authority rules. Across evaluated LLM families, separating context applicability from authority polarity improves authority-resolution accuracy. Combining this separation with priority-ordered evaluation achieves a Pareto-optimal trade-off between accuracy and token usage. Integration with AgentDojo further shows that structured resolution preserves authorized actions while reducing attack success. These findings highlight structured authority reasoning as a basis for reliable action control in multi-user environments.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.