DreamSPI: World Models for Safe Policy Improvement through Latent Imagination
Abstract
In model-based reinforcement learning (MBRL), world models promise to support policy improvement with fewer environment interactions. Safe policy improvement (SPI) can reason about such updates, but its guarantees are local to the behavior policy and difficult to combine with changing encoders, deep representations, and replay-heavy model learning. We introduce DreamSPI, a principled on-policy MBRL algorithm that translates SPI locality into practical update proxies while updating the encoder, model, and latent actor-critic. It tokenizes spatial encoder outputs, predicts categorical next-token embeddings, imagines fixed-horizon latent rollouts from fresh vectorized data, and stabilizes coupled updates with Kronecker-factored optimization. On ALE-57 and Octax, DreamSPI outperforms strong baselines while ablations identify on-policy latent imagination and our token-embedding interface as the key design.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.