Learning Recurrent Belief Representations for Diffusion Policies with Predictive World Modeling
Abstract
Diffusion policies are commonly conditioned on a fixed observation window, restricting their effectiveness under partial observability, occlusion, and tasks requiring long-horizon memory. In a partially observable Markov decision process (POMDP), an effective policy should maintain a belief state that integrates historical information into a compact, structured world representation. We propose Belief-Inferring Recurrent Diffusion Policy (\ours), which integrates recurrence with attention mechanisms to maintain a latent belief representation that conditions the diffusion policy. In addition to the imitation objective, the belief encoder is trained with an auxiliary world-modeling loss that predicts latent representations of future observations from the current belief and a future action sequence, encouraging the latent state to capture the underlying world state. Across six simulated 2D manipulation tasks, achieves a mean peak success rate of 74%, a improvement over the strongest baseline, and outperforms diffusion policies that lack a belief state or are trained solely with the imitation objective.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.