Risk Budgets for Supervising Agents with Hidden Goals under Temporal Logic Constraints
Abstract
Autonomous agents are often deployed with fixed goal-conditioned skills while their operating requirements continue to change. A planetary rover, for example, may face new ordering or deadline requirements after launch, when retraining and recertifying its flight-qualified controller is infeasible. We study an external supervisor for such post-deployment requirements, expressed in linear temporal logic on finite traces (LTLf). It observes physical states, infers the agent's active goal from motion, and may reassign that goal, while the low-level controller remains inaccessible. The central difficulty is that decisions that appear acceptable in isolation can accumulate into an episode-level violation probability above the requested tolerance. We introduce an intent-aware supervisor whose risk account prices autonomy relative to a task-level takeover policy, charging increases in continuation risk and refunding decreases. Goal-conditioned dynamic programming precomputes the violation risks of autonomy until the next task event, one autonomous step, and takeover, while Bayesian filtering combines them under the current intent belief. We prove an exact accounting identity: the bad-prefix violation probability plus the expected terminal balance equals the requested tolerance. Experiments show that the method enforces instance-specific tolerances while retaining sparse high-level intervention across various LTLf specifications.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.