Breaking the Staleness Phase: Understanding and Accelerating JEPA Training
Abstract
The exponential moving average (EMA) mechanism, introduced to stabilize self-supervised learning, coincides with a surprising failure mode in Joint-Embedding Predictive Architectures (JEPAs): during early training, the reconstruction loss plummets while learned representations remain useless for downstream tasks. We call this the “staleness” phase. It arises from two interacting effects: the EMA keeps the target encoder tightly coupled to the main encoder in parameter space, and the predictor, which trains much faster than the encoder, fits the targets of the still uninformative encoder, minimizing the objective while starving the encoder of gradient. We derive an upper bound showing that the coupling depends on the interplay between momentum dynamics and augmentation strategy, and that escape requires the encoder's updates to outpace the EMA's smoothing. This reframes staleness not as an objective flaw but as a transient regime with predictable onset and recovery. The analysis suggests a simple fix: perturbing encoder weights at the start of training breaks the initial symmetry between encoder and target, while the perturbation is quickly absorbed by the EMA and vanishes as training matures. Experiments on images and time-series support the analysis and show that noise injection accelerates recovery and improves final performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.