Safe Online Personalization of Agent Permissions via Auditable Policy Distillation
Abstract
Personal AI assistants act with broad authority (reading files, sending email, executing code), while the permission policies that constrain them are written down in advance and fit no individual user. We formulate safe selective permission learning: an agent's governance layer must decide ALLOW/DENY/ASK per action online, learning each user's latent preferences from feedback that is costly to elicit and missing-not-at-random, while false allows on risky actions are budgeted per risk tier and every learned behavior remains inside immutable, user-declared policy ceilings. Our method couples a cost-sensitive Bayesian logistic learner (with ASK as an explicit observation action, inverse-propensity-corrected updates, and confidence-gated autonomy on high-risk tiers) with auditable policy distillation: learned decisions are periodically compiled into symbolic rules in a small policy grammar, admitted only after calibration-set validation, and enforced thereafter as inspectable, revocable policy. We instantiate the method as GUARD-OPENCLAW, a governance layer for the open-source OPENCLAW platform in which a deterministic trusted core validates typed, single-use, provenance-carrying capabilities at the tool boundary and writes hash-chained receipts ahead of execution. In a behavioral evaluation on the user-client side with template-grouped splits and six user personas, the learner reaches selective accuracy at autonomous coverage with a false-allow rate among autonomous allows, versus for linear Thompson sampling and for the best fixed policy (mined attribute rules), at a median end-to-end governance overhead of . The extension, benchmark generator, and a CPU-only harness that reproduces the learning dynamics of our method are released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.