TS-Dreamer: Aligning Representations in World Models at Different Temporal Scales
Abstract
World models let agents learn from imagined trajectories, so their latent states must preserve information for prediction and control. Aligning states with image embeddings provides supervision without pixel reconstruction. Single-step alignment directly supervises correspondence with the current image, but not explicitly over multiple observations, where motion and other temporal properties can become apparent. Different observation intervals can reveal longer-term patterns. We ask whether alignment at different temporal scales can provide complementary supervision. TS-Dreamer maps latent states into the image embedding space and jointly aligns the two representations at individual steps and over temporal windows of different lengths. At each scale, normalized sums over matching windows are aligned using Barlow Twins, without adding trainable parameters or changing inference. TS-Dreamer leads the evaluated methods in mean and median returns on DeepMind Control Suite and remains competitive on Atari 100K. In FlappyBird, adding alignment windows up to 32 decisions yields the two-scale score at fixed capacity and training budget. Representation analyses show improved linear readout of physical quantities and their changes. These results highlight temporal scale as a design choice for world model supervision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.