acceptodds
Under review as a conference paper at ICLR 2027

Representation Selection in Capacity-Limited LeJEPA

Abstract

Self-supervised Joint-Embedding Predictive Architectures (JEPAs) learn from correlated observations of the same system. Existing LeJEPA theory shows that, when the representation dimension matches the number of latent variables and all latent variables evolve at the same rate, the training objective can linearly recover the latent variables. In practical models, however, the representation dimension is often smaller, and there is still no complete theory of how limited representation capacity is allocated. We study latent variable selection in LeJEPA under limited representation capacity and analyze how its distributional constraint affects this selection. We show that, when latent variables evolve at different rates, the model ranks both the latent variables and nonlinear functions of them. As a result, functions of higher order of slower latent variables can replace faster latent variables in the learned representation. LeJEPA's distributional constraint can rule out simple polynomial solutions of higher order, but it can still allow multiple representation coordinates to depend on the same slow latent variable. This source reuse can leave predictable information about the next representation outside the selected span. Based on this analysis, we introduce an additional conditional mean constraint to reduce this information loss. Under the Gaussian model, the two constraints yield linear invariant structure, and alignment selects the leading linear subspace when the cutoff is strict. Our experiments support the predicted selection boundary and the effect of the proposed correction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.