Lower Loss by Forgetting: How Trainable Targets Suppress Information in JEPA World Models
Abstract
Joint-embedding predictive architectures (JEPAs) learn world models by predicting future representations and can reduce prediction error by suppressing information already present in their representations. We identify gradients through the learned prediction target as a mechanism for this suppression. Our gradient decomposition shows that, with an optimal predictor and a linear encoder aligned with independent factors, the target contribution attenuates each factor at a rate proportional to its innovation variance. A quadratic covariance penalty yields a closed-form retention threshold and links retained variance to each factor's weight in the linear model's planning distance. Controlled experiments with initially decodable synthetic signals support the predicted ordering, with higher-innovation signals losing decodability earlier. In a simulated pushing task, a model acquires and subsequently suppresses information about block position. A state variable can also remain decodable while having little influence on the latent distance used for planning. On the Cube manipulation task from OGBench, weaker regularisation reduces normalised prediction error by while lowering contact-related decodability and planning success. Paired head interventions improve Cube planning, but their benefits depend on the task. These findings identify a mechanism by which predictive training can impair a representation for control, and motivate evaluating information retention and its influence on planning distance alongside prediction accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.