Learning World Models with Contingent Self-Prediction
Abstract
World models supervised via latent self-prediction naturally prioritize predictable environment dynamics. However, complex settings are rife with predictable distractors, such as periodic motion or moving backgrounds, that are trivial to predict yet uninfluenced by an agent's actions. Crucially, standard non-collapse regularizers cannot eliminate them, as representations saturated with task-irrelevant noise can satisfy distribution constraints at zero prediction loss. To resolve this, we exploit time homogeneity: identifying features over a future planning horizon already determined by present history is equivalent to identifying features of the present that were already foreseeable from an older context. Conditioning on this older context allows us to decompose state predictions into a predetermined component (fixed by past history) and a contingent component (driven by recent actions and observations). We call the resulting approach contingent self-prediction: an objective that enforces non-collapse constraints exclusively on the contingent subspace. Our method is provably invariant to predictable distractors, ensuring state capacity is prioritized for plan-dependent dynamics.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.