acceptodds
Under review as a conference paper at ICLR 2027

Learning Recurrent Belief Representations for Diffusion Policies with Predictive World Modeling

Abstract

Diffusion policies are commonly conditioned on a fixed observation window, restricting their effectiveness under partial observability, occlusion, and tasks requiring long-horizon memory. In a partially observable Markov decision process (POMDP), an effective policy should maintain a belief state that integrates historical information into a compact, structured world representation. We propose Belief-Inferring Recurrent Diffusion Policy (\ours), which integrates recurrence with attention mechanisms to maintain a latent belief representation that conditions the diffusion policy. In addition to the imitation objective, the belief encoder is trained with an auxiliary world-modeling loss that predicts latent representations of future observations from the current belief and a future action sequence, encouraging the latent state to capture the underlying world state. Across six simulated 2D manipulation tasks, achieves a mean peak success rate of 74%, a improvement over the strongest baseline, and outperforms diffusion policies that lack a belief state or are trained solely with the imitation objective.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.