WORLDPATCH: RETAINED PREDICTIVE STATES FOR ANTICIPATORY MULTIMODAL DECODING
Abstract
Predictive representation learning can use future observations to shape visual representations, yet the predicted state is often discarded before downstream decoding. Consequently, gains from future supervision do not reveal whether predictive training merely improves the representation of the present or whether the predicted state itself is useful when retained. We introduce WorldPatch, which predicts a short-horizon latent successor from the current image and retains it alongside the present representation for decoding. Paired future images supervise the successor during training but are unavailable at inference. We vary decoder access under matched predictive supervision to distinguish the effects of retaining the predicted successor from those of predictive training. On EPIC-KITCHENS-100, WorldPatch improves anticipatory decoding over an auxiliary-only model, with consistent gains on a second backbone and on Ego4D. Further analysis on the primary backbone shows that retaining the successor yields larger gains when current and future action labels differ. These findings show that future prediction can be useful beyond its role as a training objective: retaining its output can improve short-horizon multimodal decoding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.