Offline Distributional World Model Value Iteration
Abstract
Learned world models enable scalable planning in reinforcement learning but remain challenging in the offline setting, where planning can exploit erroneous rewards and transitions induced by model approximation error. While prior work has established distributional value iteration under known dynamics, little is known about distributional planning over learned world models in partially observable environments. We address this gap by deriving Wasserstein bounds that characterize how world-model approximation error affects latent return distributions and by introducing a support-dependent reward shift with distributional convergence guarantees under stated assumptions. Building on this analysis, we propose Offline Distributional World model Value Iteration (ODWVI), a practical algorithm for support-constrained distributional planning in learned latent world models. Across three offline planning benchmarks, ODWVI achieves the highest mean return among the evaluated methods in the main benchmark configurations. Support ablations further show that performance degrades as planning is extended to increasingly weakly supported actions, highlighting the importance of behavior-policy support when planning with imperfect learned world models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.