acceptodds
Under review as a conference paper at ICLR 2027

AuthDiff: Measuring and Constraining Authorization Drift in LLM Agents

Abstract

LLM agents update their authorization policies during execution to accommodate task requirements revealed by tool results. Existing agent-security evaluations focus on attack success rate and task completion, leaving changes to the policies themselves unmeasured. We identify and define authorization drift: successive policy updates permit requests that were initially denied, without a supporting trusted grant. Drift can make the monitor permit later messages, file changes or transfers beyond the initial authorization boundary, even when tasks succeed and injections fail. We introduce AuthDiff to measure and constrain drift. AuthDiff compares decisions on the same requests before and after each update, checks expansions against the initial policy and trusted grants, and uses an online guard to reject flagged updates. We observe drift in two dynamic authorization systems, Progent and DRIFT. All eleven evaluated agent models exhibit drift under Progent. On GPT-5.6-Luna, the guard reduces drift events per update by 86.9% while recording the same number of successful tasks. Our findings show that constraining authorization drift is essential to securing LLM agents throughout task execution.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.