acceptodds
Under review as a conference paper at ICLR 2027

SphereJEPA: Learning Spherical Video Representations From Perspective Video via Self-Supervision

Abstract

We introduce SphereJEPA, a joint-embedding predictive architecture that learns to predict features for every direction around the camera from ordinary perspective video. Its context and target come from different projections of the same scene. The context encoder reads the perspective video, an EMA target encoder reads the panorama, and both write onto tokens indexed by direction on a gravity-aligned sphere. Unlike a standard JEPA, whose predictor is discarded after pre-training, SphereJEPA keeps its predictor, because its output is the representation of what lies out of view. Training crops are rendered from the panorama, so the data needs no posed or calibrated capture, and the model never receives camera parameters. Inference is a single forward pass through a 113M-parameter encoder and predictor. The predicted features identify unseen scenes and their orientation, and agree with what the camera later observes. With frozen features and an analytical readout, SphereJEPA roughly halves the TAPVid360 mean angular error of every method not trained on the benchmark, without using point labels. representation whose spatial support extends beyond the current image, while fine-grained transfer depends on the complete training construction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.