Recurrent Structure Reveals How Goal-Conditioned Policies Fail
Abstract
Failed trajectories may follow different paths yet converge to the same persistent outcome. We propose a behavioural interpretation of goal-conditioned policy failure whose unit is a recurrent outcome of the frozen policy–environment dynamics, its incoming basin, and its operating probability mass. This identifies which starts share a failure, where behaviour becomes trapped, and how consequential that failure is. In finite deterministic systems, the closed-loop map yields an absorbing goal outcome or non-goal fixed points and cycles. Across FourRooms and MiniGrid KeyDoor, this representation reveals failure concentration and task-stage bottlenecks. In the larger 16 × 16 KeyDoor setting (MG16), exact analysis identifies 1,755 goal-conditioned failure structures across 84 right-room goals. Discovery from 32 starts per goal—8.2% of the evaluated start population—recovers 85% of pooled recurrent failure mass and identifies the dominant failure in 98.8% of goal–trial pairs. With 64 starts, mean per-goal coverage reaches 91.3%. Discovery uses sampled executions; exact enumeration supplies evaluation ground truth. The representation also guides intervention. Sample-selected FourRooms edits redirect incoming basins toward success, while iterative MiniGrid repair addresses downstream failures. Recurrence-guided exploration and replay improve terminal right-goal success over ordinary continuation in all three tested MG16 seeds, with heterogeneous gains. These results support the recurrent outcome, basin, and mass as a recoverable and actionable unit of behavioural interpretation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.