A Plug-and-Play Runtime Safety Layer for Generalist Robot Policies
Abstract
Vision-language-action (VLA) policies and world-action models (WAMs) provide increasingly general interfaces for robotic manipulation, yet their action genera- tion is primarily optimized for task completion rather than explicit safety-critical execution. As a result, semantically reasonable actions may still be physically unsafe, leading to inaccurate grasps, collisions, or unstable placements and ex- posing a gap between the general capabilities of modern robot policies and the execution-level safety requirements of manipulation. We introduce a plug-and-play, training-free runtime safety layer that operates between a frozen generalist policy and the robot controller. Rather than replacing the underlying policy, the proposed framework selectively intervenes only at safety-critical stages and uses specialized runtime modules to detect and handle unsafe actions, while leaving normal pol- icy execution unchanged. This design requires neither policy fine-tuning nor an additional vision-language-model planner, allowing the original policy to retain its general-purpose capabilities while gaining more reliable execution at critical manipulation stages. We evaluate the framework on manipulation tasks involving target ambiguity, contact-critical grasping, transport hazards, and unsafe placement using multiple frozen backbones. Results show that the proposed runtime safety layer improves both task success rate and safe success rate, while its plug-and-play design enables the same framework to be applied across different generalist robot policies without modifying or retraining the underlying models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.