acceptodds
Under review as a conference paper at ICLR 2027

Sanctuary: Safe Correlated Exploration for Online Vision-Language-Action Models Adaptation

Abstract

Online reinforcement learning (RL) can adapt vision–language–action (VLA) models to new tasks, but exploration during RL exposes robots to physical risks. Execution safeguards can prevent hazardous actions from being carried out, but the policy still needs to learn to avoid proposing them. We introduce SANCTU- ARY, a framework that preserves this risk information to guide exploration under real-time execution protection. SANCTUARY evaluates each complete action chunk before intervention and combines its risk score with the subsequent execu- tion outcome to learn long-term risk. Task value and long-term risk then jointly guide the action mean and exploration correlations across time and action dimen- sions, while marginal noise variances remain fixed. This allows blocked hazardous proposals to inform policy improvement even when protection prevents physical violations. On LIBERO-Safety and 12 custom simulation scenarios, SANCTU- ARY improves success rates by 6.3–9.0 percentage points over the strongest base- line in each evaluation, with no observed physical violations. Ablations support the complementary roles of risk learning and execution protection

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.