Aligning Video World Models with Executable Robot Actions via Inverse Dynamics Rewards
Abstract
Video generative models offer a promising basis for robotic planning by predicting future visual rollouts from observations and task instructions. However, their training objectives typically lack explicit constraints on robot executability. Generated rollouts can exhibit embodiment-related artifacts, including arm deformation and temporal discontinuities, while the actions decoded from these rollouts may be unstable or violate robot motion limits. We term this mismatch between generated robot behavior and embodiment-specific execution requirements the *executability gap*. We introduce **Executable Video Alignment (EVA)**, a post-training framework that uses action-space feedback to improve both the visual integrity of generated robot motion and the feasibility of decoded actions. EVA trains an inverse dynamics model (IDM) on robot demonstrations and freezes it to decode actions from generated videos. We compute a reward from these actions that penalizes non-smooth motion and violations of velocity and acceleration limits, and use it to guide reinforcement-learning updates to the video generator. By feeding decoded-action constraints back into video generation, EVA links visual and action-level improvements through a shared alignment signal. Experiments on RoboTwin, LIBERO, and a real bimanual robot show that EVA reduces embodiment-related visual artifacts and improves task success when executing IDM-decoded actions. Further experiments show that video models fine-tuned with EVA generate more effective video–action pairs for training downstream VLAs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.