acceptodds
Under review as a conference paper at ICLR 2027

You Only Live Once: Stop the Gradient or Regularize and Replay?

Abstract

An agent that learns a world model from its own experience sees one finite trajectory, streamed in time order. We study JEPA in this you-only-live-once (YOLO) setting and find that the standard stop-gradient target fails silently: with finite data, the encoder can lock onto spurious noise correlations while its latent variance and training loss look healthy. An analysis of the objective motivates this failure. Removing stop-gradient and adding a variance floor regularizer prevents the collapse, and replaying past experience widens the samples for estimating the variance. Our Replay Variance Floor (RVF) recovers the latent dynamics in controlled linear systems and improves planning over the stop-gradient baseline on a range of visual control tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.