DiPA-WM: Diffusion Policy Adaptation with Continual World Modeling
Abstract
World models enable sample-efficient policy learning by allowing agents to optimize action decisions through imagined rollouts. In the offline-to-online reinforcement learning setting, the data distribution often shifts as the policy improves, requiring the world model to continuously adapt to on-policy data. This creates a fundamental challenge for diffusion-based policies, whose denoising process depends on stable conditioning representations: updating the world model causes , while freezing it prevents tracking the evolving policy distribution. In this work, we propose (ffusion olicy daptation with Continual orld odeling), which addresses this tension through a factorized two-stage design. The world model is decomposed into a and a . In Stage 1, the world model is pretrained on offline data and a 1-step Gaussian policy is learned over imagination; and the context encoder is frozen after Stage 1 training. In Stage 2, a diffusion policy is initialized via behavior cloning and then trained further over imagination while the adapter is updated on freshly collected on-policy data. The frozen context encoder provides stationary conditioning features, mitigating representation drift; the world model adapter enables continual adaptation of the world model dynamics to the evolving policy distribution. To address possible non-stationarity, we further develop a nested two-loop adaptation for the world model and diffusion policy, facilitated by periodic updates of context encoder on a larger timescale. DiPA-WM is modular and general: it consistently improves multiple diffusion policy optimization methods across continuous locomotion, robotic manipulation, and autonomous driving benchmarks, outperforming frozen-world-model and model-free baselines without modifying any core training objectives.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.