VLSA-RL: Can a Small Residual Policy Enable Safety-Aware VLA Adaptation?
Abstract
How can a pretrained Vision-Language-Action (VLA) policy adapt to safety-critical manipulation tasks without updating its billions of parameters? Execution-time safety filters can prevent unsafe actions but do not teach the policy to avoid repeated interventions, while end-to-end reinforcement learning (RL) is costly. We introduce VLSA-RL, which freezes the VLA backbone and learns a 1.3M-parameter residual policy (0.037% of ) to adjust its actions before control-barrier-function (CBF) safety filtering. A safety-aware GRPO objective uses feedback from the CBF quadratic program, risk-dependent clipping, and task and safety rewards to optimize the residual policy. We further analyze how residual actions can compensate for dynamics mismatch and modulate CBF convergence behavior. On SafeLIBERO, VLSA-RL reaches a 68.0% average safe-success rate across two safety levels. Relative to the CBF-filtered AEGIS baseline, it improves safe success by 21.4 and 22.9 percentage points on Levels I and II, respectively, while reducing average episode length on both levels.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.