Closing the Loop on Masking Difficulty for Joint-Embedding Predictive Learning
Abstract
Joint-embedding predictive architectures such as I-JEPA learn visual representations by predicting the embeddings of target blocks from a visible context. The masking policy that defines this task is typically fixed (I-JEPA samples M=4 target blocks throughout training), even as the representation it shapes evolves. However, a fixed mask does not imply fixed difficulty: the same masking operation can displace the representation differently at different stages of learning. We introduce Feedback-Controlled Data Augmentation (FCDA), a closed-loop framework that adapts masking using online feedback from representation geometry. A gradient-free sensor measures masking-induced displacement along weak but reliably estimated directions of the feature covariance, and a controller adapts the number of target blocks M so that this displacement tracks a reference calibrated on the original I-JEPA baseline, leaving the training objective untouched. In a matched- seed ImageNet-1K study with ViT-Tiny/16, FCDA improves linear-probe top-1 accuracy by 3.13 percentage points over the original baseline and also outperforms fixed and approximately target-count-matched open-loop controls. Transfer gains on natural-image recognition and a broader final covariance spectrum support the approach, while best-of-four CLEVR performance remains close to baseline. These results demonstrate that FCDA learns stronger representations by treating masking as a feedback-controlled component of predictive learning, one that adapts to the representations it helps shape.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.