acceptodds
Under review as a conference paper at ICLR 2027

Characterizing How JEPA Constructs Representations

Abstract

In joint-embedding predictive architectures (JEPA), an encoder and a predictor are trained together to produce a representation that minimizes the error in predicting the representation of the next observation. Because the representations of both the current and the next observation are learned, it is unclear which features they will include. Earlier works portray JEPA as selecting features based on their individual predictability, raising the concern that control-relevant features may be omitted if they are hard to predict. To address this concern and better understand how JEPA selects which features to encode, we propose the predictive closure characterization. Based on a reformulation of the JEPA training objective, we show that selection depends on joint properties of features. We support this characterization with experiments showing that unpredictable features can become better represented when they help predict other features, and that changing a feature's predictive relationships can affect how well others are represented. We also demonstrate implications for steering feature selection: adding an auxiliary objective targeting one feature can strengthen the representation of others, and supplying a feature as side information to the predictor can change how well other features are represented. We hope this work helps practitioners understand and steer JEPA representations and guides researchers in improving how these systems learn.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.