acceptodds
Under review as a conference paper at ICLR 2027

PILOT: Self-Supervised Visual Representation Pretraining for Temporal Correspondence

Abstract

Self-distillation has emerged as a strong paradigm for self-supervised visual representation learning. However, extending this approach to learning video representations with strong temporal correspondence remains challenging: raw frames provide no explicit patch-to-patch alignment over time, while self-supervised pretraining must derive correspondence supervision without external optical flow, point tracks, or correspondence annotations. We address this bottleneck with PILOT (Privileged Information Learning via Optimal Transport), a self-supervised representation pretraining framework based on explicit cross-frame target construction. PILOT uses asymmetric distillation: a privileged teacher combines information privilege with algorithmic privilege (Sinkhorn-Knopp optimal transport) to produce globally constrained cross-frame patch assignment predictions that serve as targets for a restricted student. These targets exhibit reduced assignment imbalance and many-to-one collisions, while the student predicts them from masked inputs using row-wise Softmax, encouraging the student encoder to internalize cross-frame correspondence. After pretraining, only a standard encoder is retained. Experiments show that PILOT achieves state-of-the-art overall performance among self-supervised pretrained encoders across three dense correspondence benchmarks and demonstrates strong transfer to action recognition and video retrieval.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.