You Only Live Once: Stop the Gradient or Regularize and Replay?
Abstract
An agent that learns a world model from its own experience sees one finite trajectory, streamed in time order. We study JEPA in this you-only-live-once (YOLO) setting and find that the standard stop-gradient target fails silently: with finite data, the encoder can lock onto spurious noise correlations while its latent variance and training loss look healthy. An analysis of the objective motivates this failure. Removing stop-gradient and adding a variance floor regularizer prevents the collapse, and replaying past experience widens the samples for estimating the variance. Our Replay Variance Floor (RVF) recovers the latent dynamics in controlled linear systems and improves planning over the stop-gradient baseline on a range of visual control tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.