Generalized Pseudo-Label in Point Tracking
Abstract
Reliable trajectory supervision is fundamental to robust point tracking, yet dense point-level annotations for real-world videos remain severely limited. To alleviate this scarcity, recent studies have increasingly explored two pseudo-labeling paradigms: dedicated methods that derive pseudo-labels from unlabeled videos and self-training approaches that reuse a tracker’s own predictions as pseudo-labels. These approaches rely on different assumptions and mechanisms, resulting in distinct prediction characteristics and varying reliability across video sequences, frames and points. To account for this variability, we propose a pseudo-label generation framework consisting of a Candidate Curator and a Coordinate Recomposer. Specifically, the Candidate Curator constructs context-aware representations from local evidence around the query and candidate points. It then reasons over temporal context and frame-level candidate relations to select reliable candidates for each point and frame. However, independent frame-wise selection can introduce temporal discontinuities in the pseudo-label trajectory. We therefore introduce the Coordinate Recomposer, which models query-conditioned spatiotemporal correspondence to recompose the selected candidate into a temporally coherent point prediction. Extensive experiments demonstrate that the proposed framework produces more accurate pseudo-labels and provides effective supervision for real-world point-tracking adaptation. Fine-tuning existing point trackers with our generated pseudo-labels further improves their performance on real-world videos.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.