acceptodds
Under review as a conference paper at ICLR 2027

AhJEPA: A single branch framework that captures granular details for JEPA methods

Abstract

Regularization-based self-supervised learning methods enable stable training without additional heuristic designs. However, such methods do not encourage learning of fine-grained relationships, which limits their applicability on downstream tasks. We mathematically characterize contextual prediction from incomplete observations and derive its implementation through a shared backbone, without a separately parameterized teacher or EMA updates. Our analysis interprets stop-gradient (SG) as a restriction on target-side adaptation, and an exploratory complementary-subset objective illustrates a possible implementation without SG. Inspired by previous masked-image-modeling (MIM) methods, we propose AhJEPA, a simple yet effective architecture that is broadly Applicable to heuristic-free JEPA losses to encourage learning of fine-grained relationships of the input data while maintaining the training simplicity and stability. Similar to previous works, we add a reconstruction loss to force the model to learn the fine-grained details. Different from them, AhJEPA maintains a one-branch training paradigm and applies the reconstruction loss to CLS tokens rather than patch tokens. We retain SG in the final method for its stronger empirical performance. We find that this contextual learning approach is a more general way to improve JEPA methods with different special designs (e.g., sparsity creation). This simple method consistently improves the performance of heuristic-free JEPA regularization methods across in-domain and out-of-domain classification, segmentation, depth estimation, and latent world model planning. Specifically, it helps the VISReg method achieve the best OOD accuracy and reduce the gap between heuristic-free methods and iBOT on fine-grained tasks. Also, it boosts the LeWorldModel performance to SOTA on OGBench-Cube and Two-Room datasets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.