What Data Makes World Models Better Manipulation Planners? Beyond Successful Demonstrations with Diverse Action Experience
Abstract
World models equip robots with the ability to predict the physical consequences of candidate actions prior to execution. However, the interaction data used to train these models is typically dominated by successful expert demonstrations. While such data ensures broad task-relevant state coverage, it offers limited experience of how alternative actions from similar states yield distinct future outcomes. In this work, we systematically investigate how the composition of robot interaction data influences world-model planning. We introduce Diverse Action Experience (DAE), a data-collection framework that branches alternative action trajectories from intermediate states of expert demonstrations, systematically allocating the interaction budget between broad state coverage and dense local action-outcome coverage. Evaluating multiple world-model architectures on the LIBERO benchmark, we measure performance through candidate-action ranking and closed-loop manipulation. Our experiments reveal that data composition substantially impacts planning behavior: varying state coverage, perturbation regimes, and interaction durations produces substantial differences in ranking fidelity and downstream task success. Furthermore, the preferred data composition depends strongly on the underlying world model. We also show that high aggregate ranking accuracy can obscure degenerate candidate selection, and that stronger candidate ranking does not automatically guarantee robust closed-loop control. These findings establish interaction-data composition as an important design dimension in world-model planning, alongside model architecture and planner design.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.