acceptodds
Under review as a conference paper at ICLR 2027

Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video–Action Modeling

Abstract

Video generation models encode rich spatiotemporal priors about scene dynamics. To serve robot control, a general video-action model must reason bidirectionally: predicting future observations under given actions, and inferring actions from observed or desired outcomes. Existing approaches fall short—either predicting joint-space vectors through dedicated action heads that vary across embodiments, or visualizing only end-effector quantities that discard full-arm articulation and still require embodiment-specific inverse-kinematics mappings. We present Dream4ACT, a video-action world model that renders target joint configurations as four action views via URDF-based forward kinematics from prescribed virtual cameras independent of physical observation cameras. These visual sequences preserve embodiment-specific geometry while sharing a video-modeling pipeline with RGB observations. A diffusion transformer jointly processes multiview physical-camera observations and four action-view streams, conditioned on semantic features from a frozen vision-language model followed by a trainable adapter. Masked flow matching supports forward dynamics, inverse dynamics, and joint observation-action generation by varying which future sequences are corrupted. For execution, training-free multiview matching against URDF-based renderings recovers joint targets without a learned embodiment-specific decoder. Dream4ACT achieves 88.98% success rate on RoboTwin 2.0, and 65.66 score on TriWorldBench. Ablation studies further confirm the benefit of action-view conditioning for high-quality video generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.