Bellman-Corrected Dualization for Safe Vision-Language-Action Models
Abstract
Vision-language-action (VLA) models offer a promising framework for generalist robot policies, but their deployment in long-horizon embodied environments faces critical safety challenges. Standard post-training approaches struggle to maintain a stable reward–safety trade-off under policy-induced distribution shifts, as evolving policies alter trajectory occupancy and invalidate static constrained alignment. We propose Dualized-SafeVLA, a Bellman-corrected dualization framework that constructs a local one-shot dual under frozen occupancy, producing a low-dimensional dual, a closed-form target policy, and trust-region-safe updates. This approach systematically aligns VLA with safety constraints while maintaining task performance under evolving distributions. On safety-critical long-horizon mobile manipulation benchmarks, Dualized-SafeVLA improves task success by 9.04 points and reduces cumulative safety cost by 79.9% compared to a clean-environment SafeVLA baseline under matched-step evaluation, and further achieves 6.03 points SR improvement with 93.3% cost reduction under a 2.1M-step budget, demonstrating a practical and robust path toward safer VLA post-training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.