Structure Beyond Pixels: Aligning Video with Structural Preferences for Robust Policy Generation
Abstract
Video generation models can serve as world models for robotic policy generation. However, because they are trained to produce visually plausible videos without physical constraints, the generated rollouts often exhibit an executability gap, leading to downstream task failures. To address this, we propose Structure-Aligned Direct Preference Optimization (Struct-DPO), which aligns a pre-trained video generator with structural preferences derived from environment interaction. Specifically, we treat physically grounded simulator observations as winning samples and model rollouts as losing samples, automatically identifying the most severe violations via structural metrics for viewpoint and geometric consistency. Notably, this approach requires neither architectural changes nor explicit 3D supervision; in our experiments, it trains only 3.95% of the base model parameters at 0.27% of the additional step of standard supervised fine-tuning, yielding high parameter and computational efficiency. Experiments on the RoboCasa benchmark across 24 manipulation tasks show consistent gains in task success, and alignment learned from a subset of tasks transfers to unseen tasks, indicating that the distilled physical priors generalize across tasks and environments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.