ForeAlign: Anticipatory Self-Alignment of World Action Models for Safer Manipulation
Abstract
World Action Models (WAMs) advance robotic manipulation by jointly predicting actions and their consequences, but greater autonomy also expands the range of potential hazards. Existing safety approaches rely largely on predefined specifications and costly environmental interaction, limiting safety coverage and making risk discovery expensive. We propose ForeAlign, a framework for Anticipatory Self-Alignment that iteratively probes imagined futures, incorporates safety feedback into policy updates for continuous improvement. It unifies heterogeneous safety requirements under a general safety distribution and extends Diffusion Negative-aware Fine-Tuning (DiffusionNFT) into a monotonic improvement operator with theoretically bounded policy shift, mitigating the safety-performance trade-off from catastrophic forgetting. To counter alignment faking, ForeAlign decouples the WAM into a cascade of an action policy and an independently parameterized action-conditioned world model, directing safety supervision to action generation. Across simulated and real-world embodied manipulation benchmarks, ForeAlign improves safe rate by 35.6% and task success rate by 24.6% on average.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.