Predicting Irreversible Transitions for Safe Switching in Autonomous RL
Abstract
Autonomous reinforcement learning (ARL) aims to reduce manual resets by alternating between forward and reset policies. One of the main challenges in ARL is handling irreversible states, from which the agent cannot recover without external intervention. While several recent works have proposed ARL algorithms that can handle irreversible states, these algorithms require task-specific knowledge, such as privileged state information or reversibility labels. Furthermore, they pay little attention to safely switching between forward and reset policies. In this paper, we propose a novel ARL algorithm that identifies irreversible states from image observations without task-specific knowledge and enables safe policy switching. Our algorithm introduces an irreversibility estimator that takes the current image observation and action as input and estimates the probability of transitioning to an irreversible state. We train the estimator on the state-action sequences that lead to irreversible states, so that it predicts an irreversible transition several steps before it occurs and aborts the ongoing policy early. The supervisory signals for these sequences are extracted from rollouts using vision-language models. Our algorithm also introduces a transition policy that guides the agent to safe switching states. These states are reversible, not too difficult for resuming the forward policy, and reachable from the aborted state. Experimental results demonstrate that our algorithm outperforms prior ARL algorithms across diverse navigation and manipulation tasks, even without task-specific knowledge.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.