TRIP: Self-Supervised Temporal Representations for Imagination and Planning
Abstract
Visual world models often use frame-wise image representations as prediction targets and to construct planning costs, without explicitly incorporating temporal context into the representations themselves. We introduce TRIP (Temporal Representations for Imagination and Planning), a framework that adapts an image-pretrained visual encoder through causal video self-supervised learning and uses the resulting representations for both future prediction and visual planning. For prediction, a compact bottleneck adapter learns to match the temporal encoder's prototype distributions, providing low-dimensional targets for a diffusion-based feature predictor whose outputs condition RGB video generation. For planning, we reuse the frozen temporal encoder to score imagined rollouts by comparing the goal's prototype distributions with and without rollout context. The resulting goal-representation consistency cost enables model predictive control without additional cost-specific training. TRIP outperforms Diffusion Forcing in video generation on SSv2 and LIBERO and achieves higher average closed-loop planning success than AdaWorld on Procgen. Its scoring encoder also transfers from Procgen to VP without further adaptation, improving planning when paired with an existing world model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.