ASR-VLA: Adaptive Safety Reasoning with Safety-Critical Future Prediction for Vision-Language-Action Models
Abstract
Vision-Language-Action (VLA) models have demonstrated strong capabilities in language-conditioned robotic manipulation. However, safety remains a critical and insufficiently addressed challenge for VLA models. Existing safety-aware VLA methods primarily formulate safety as a policy constraint or apply external correction to potentially unsafe actions. They generally do not explicitly ground action generation in a predicted, state-dependent future hazard and reason about how the robot should respond. Consequently, the policy may struggle to determine an appropriate response to evolving risks while maintaining task progress. To address this problem, we introduce ASR-VLA, an adaptive framework that integrates explicit safety reasoning into VLA policies. ASR-VLA couples a risk-aware world model with a vision-language reasoning model. To support this reasoning, ASR-VLA incorporates a risk-aware world model that predicts the most safety-critical future observation as auxiliary visual evidence. The vision-language reasoning model analyzes the current scene and task context, using the predicted future risk to generate structured safety guidance. The resulting reasoning and future visual evidence jointly condition action generation, enabling the policy to adapt its behavior while maintaining task progress. Experiments on SafeLIBERO demonstrate that ASR-VLA substantially improves both task performance and safety, achieving a 93.68% success rate with only a 9.46% collision rate. Compared with the strongest baseline, ASR-VLA increases task success by 20.72% while reducing collisions by 10.46%, yielding a markedly better success–safety performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.