Next Embedding Prediction Makes World Models Stronger
Abstract
Dreamer-style world models process trajectories sequentially, yet their representation losses are usually anchored to the current observation through pixel reconstruction or same-timestep latent alignment. This mismatch is limiting under partial observability, where successful control depends on retaining compact information about cues, landmarks, and spatial context long after they disappear. We introduce NE-Dreamer, a decoder-free world model that moves the representation target one step into the future. Given the history-conditioned latent state, a lightweight causal predictor forecasts the next encoder embedding, and training aligns this prediction to a stop-gradient target with a redundancy-reduction objective. This converts representation learning from same-timestep description into history-conditioned prediction, directly pressuring the latent state to preserve information useful beyond the current frame. Across bsuite diagnostics, MiniGrid Memory, and DMLab, NE-Dreamer improves most strongly in long-delay memory and navigation settings, while matching strong decoder-based and decoder-free baselines on the DeepMind Control Suite. These results identify temporal target placement as a distinct and complementary lever to architectural memory for improving partially observable model-based reinforcement learning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.