acceptodds
Under review as a conference paper at ICLR 2027

TRIP: Self-Supervised Temporal Representations for Imagination and Planning

Abstract

Visual world models often use frame-wise image representations as prediction targets and to construct planning costs, without explicitly incorporating temporal context into the representations themselves. We introduce TRIP (Temporal Representations for Imagination and Planning), a framework that adapts an image-pretrained visual encoder through causal video self-supervised learning and uses the resulting representations for both future prediction and visual planning. For prediction, a compact bottleneck adapter learns to match the temporal encoder's prototype distributions, providing low-dimensional targets for a diffusion-based feature predictor whose outputs condition RGB video generation. For planning, we reuse the frozen temporal encoder to score imagined rollouts by comparing the goal's prototype distributions with and without rollout context. The resulting goal-representation consistency cost enables model predictive control without additional cost-specific training. TRIP outperforms Diffusion Forcing in video generation on SSv2 and LIBERO and achieves higher average closed-loop planning success than AdaWorld on Procgen. Its scoring encoder also transfers from Procgen to VP without further adaptation, improving planning when paired with an existing world model.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.