acceptodds
Under review as a conference paper at ICLR 2027

DreamSPI: World Models for Safe Policy Improvement through Latent Imagination

Abstract

In model-based reinforcement learning (MBRL), world models promise to support policy improvement with fewer environment interactions. Safe policy improvement (SPI) can reason about such updates, but its guarantees are local to the behavior policy and difficult to combine with changing encoders, deep representations, and replay-heavy model learning. We introduce DreamSPI, a principled on-policy MBRL algorithm that translates SPI locality into practical update proxies while updating the encoder, model, and latent actor-critic. It tokenizes spatial encoder outputs, predicts categorical next-token embeddings, imagines fixed-horizon latent rollouts from fresh vectorized data, and stabilizes coupled updates with Kronecker-factored optimization. On ALE-57 and Octax, DreamSPI outperforms strong baselines while ablations identify on-policy latent imagination and our token-embedding interface as the key design.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.