Sanctuary: Safe Correlated Exploration for Online Vision-Language-Action Models Adaptation
Abstract
Online reinforcement learning (RL) can adapt vision–language–action (VLA) models to new tasks, but exploration during RL exposes robots to physical risks. Execution safeguards can prevent hazardous actions from being carried out, but the policy still needs to learn to avoid proposing them. We introduce SANCTU- ARY, a framework that preserves this risk information to guide exploration under real-time execution protection. SANCTUARY evaluates each complete action chunk before intervention and combines its risk score with the subsequent execu- tion outcome to learn long-term risk. Task value and long-term risk then jointly guide the action mean and exploration correlations across time and action dimen- sions, while marginal noise variances remain fixed. This allows blocked hazardous proposals to inform policy improvement even when protection prevents physical violations. On LIBERO-Safety and 12 custom simulation scenarios, SANCTU- ARY improves success rates by 6.3–9.0 percentage points over the strongest base- line in each evaluation, with no observed physical violations. Ablations support the complementary roles of risk learning and execution protection
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.