Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video–Action Modeling
Abstract
Video generation models encode rich spatiotemporal priors about scene dynamics. To serve robot control, a general video-action model must reason bidirectionally: predicting future observations under given actions, and inferring actions from observed or desired outcomes. Existing approaches fall short—either predicting joint-space vectors through dedicated action heads that vary across embodiments, or visualizing only end-effector quantities that discard full-arm articulation and still require embodiment-specific inverse-kinematics mappings. We present Dream4ACT, a video-action world model that renders target joint configurations as four action views via URDF-based forward kinematics from prescribed virtual cameras independent of physical observation cameras. These visual sequences preserve embodiment-specific geometry while sharing a video-modeling pipeline with RGB observations. A diffusion transformer jointly processes multiview physical-camera observations and four action-view streams, conditioned on semantic features from a frozen vision-language model followed by a trainable adapter. Masked flow matching supports forward dynamics, inverse dynamics, and joint observation-action generation by varying which future sequences are corrupted. For execution, training-free multiview matching against URDF-based renderings recovers joint targets without a learned embodiment-specific decoder. Dream4ACT achieves 88.98% success rate on RoboTwin 2.0, and 65.66 score on TriWorldBench. Ablation studies further confirm the benefit of action-view conditioning for high-quality video generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.