DELE-w0.5: Inferring Action from Future Latent State for Robotic Manipulation
Abstract
Video-generation-based World-Action Models (WAMs) jointly predict dense future visual trajectories and robot actions. For robotic manipulation, we investigate whether action learning can benefit from predicting physical outcomes without reconstructing every intermediate frame. In this paper, we propose DELE-w0.5, which learns robot actions jointly with future latent-state prediction, without video generation. The future latent state models action-relevant physical outcomes and connects world modeling with action learning through shared representations. The principle of DELE-w0.5 is to model how the real world changes under robot actions, rather than how its visual appearance evolves frame by frame. Across 640 real-robot trials on four long-horizon manipulation tasks, DELE-w0.5 achieves the best performance among all compared policies, attaining 62.5% overall full-task success and 81.3% macro ordered-stage progress. It outperforms the strongest baseline by 32.5 percentage points in full-task success and 20.1 percentage points in macro progress.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.