DSSF: Robust Latent State Construction for JEPA-Based Visual Planning
Abstract
Latent world models based on the Joint-Embedding Predictive Architecture provide an efficient approach to dynamics modeling for visual planning by predicting future representations rather than reconstructing pixels. However, latent prediction primarily optimizes the predictability of visual information and does not guarantee that the retained factors are relevant to control. Consequently, task-irrelevant backgrounds and visual distractors may be encoded into the latent state, perturbing state prediction and goal matching when previously unseen visual distractors occur at test time. We propose DSSF, a two-stage latent-state construction method for improving planning robustness under unseen visual distractors. In the first stage, learnable prototypes model the feature-space support of non-distractor visual factors and filter out local factors that are not supported by the training distribution. In the second stage, DSSF sparsely selects information relevant to action-conditioned dynamics from the retained factors to construct a compact latent state. In this way, DSSF separately models whether a visual factor is supported by the training distribution and whether it is relevant to control, thereby suppressing both unseen distractors and in-distribution yet control-irrelevant information. We evaluate unseen distractor instances and cross-family distractors to systematically assess planning robustness under increasingly severe shifts in distractor appearance and count. Across multiple continuous-control tasks, DSSF better preserves clean-setting planning performance and consistently outperforms strong baselines under unseen distractors and cross-family distractor-appearance shifts, indicating that its robustness does not rely on memorizing the appearances of training distractors.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.