RBGA: Residual-Budget Guide Arbitration for Offline-to-Online Safe Reinforcement Learning in Dexterous Manipulation
Abstract
Dexterous manipulation requires robots to leverage existing datasets for rapid skill acquisition while maintaining safety during real-world deployment to prevent damage to the robot or its surroundings. Offline-to-online safe reinforcement learning (O2O safe RL) provides a promising solution for this problem by leveraging prior experience for online improvement under safety constraints, thereby mitigating the sample inefficiency and exploration risk of learning safe policies from scratch. However, existing O2O safe RL approaches do not explicitly account for the safety budget already consumed when making decisions within an episode. Excessive cost incurred early can therefore leave little room for subsequent exploration and policy improvement. To address this limitation, we propose Residual-Budget Guide Arbitration (RBGA), an algorithm that explicitly tracks the remaining episode safety budget and uses it to adaptively select between offline guidance and the evolving online policy. At each step, RBGA compares the cost-to-go estimates of the guide and online actions against the remaining budget, enabling step-wise action arbitration throughout the episode. We provide theoretical analysis characterizing the safety implications of residual-budget-based selection and its effect on guide-assisted exploration. Extensive experiments demonstrate that RBGA enables effective online policy improvement while maintaining strong safety performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.