Learning Successor-Measure World Models for Zero-Shot Control from Pixels
Abstract
World models typically support decision-making by predicting future observations or latent states, with behavior adapted to a task through planning or policy optimization. For zero-shot control, we instead ask whether a world model can directly organize the long-horizon visitation induced by different behaviors, so that a new reward can select among them without additional policy optimization. We propose SWM, a successor-measure world model for zero-shot visual control. Rather than learning step-by-step predictive dynamics, SWM learns policy-conditioned future occupancy in a compact pretrained visual latent space. SWM uses a frozen DINO encoder as a task-agnostic visual basis and trains a lightweight adaptor, together with forward and backward successor-measure representations, to organize visual observations into a space where future occupancy can be factorized and reused across tasks. At test time, a new task specified by reward-labeled observations is projected onto the learned backward basis, directly selecting a policy from the pretrained behavior family without additional policy optimization or online interaction. Experiments on pixel-based Control tasks show that SWM consistently improves over reward-free predictive world-model baselines and achieves higher mean return than CNN-FB on all sixteen evaluated tasks. Importantly, the learned latent space preserves control-relevant pose information and retains robustness under visual perturbations, suggesting that pretrained visual features can serve as a reusable basis for occupancy-based world modeling when aligned through successor-measure factorization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.