SyncTAP: Tracking Any Point from Unsynchronized Multi-View Videos
Abstract
Multi-view 3D point tracking is an important prior to understand dynamic 3D scenes, but typically assumes temporally aligned input videos. Even small time offsets collapse the underlying epipolar geometry, disrupting cross-view feature aggregation and severely degrading tracking performance. We propose SyncTAP, the first unified tracking framework that jointly estimates sub-frame per-view time offsets and high-fidelity one-to-one correspondent 2D/3D trajectories from unsynchronized multi-view videos. First, our Epipolar-based Time Offset Initialization (EpiInit) analytically finds initial time offsets using an epipolar constraint and a 2nd-order Taylor approximation. Then, a transformer-based refinement stage alternates two operations: (i) an Epipolar-based Time Offset Refinement (EpiRef) module updates time offsets via Sampson error landscapes; (ii) a Spatio-Temporal-Cross-View (STCV) attention module refines 2D/3D trajectories by introducing novel attention layers for robust feature aggregation under temporal misalignment. This joint optimization yields a mutually beneficial cycle between synchronization accuracy and tracking accuracy. Extensive experiments demonstrate that our SyncTAP achieves sub-frame synchronization accuracy, significantly outperforming SOTA visual-based synchronization method and establishing a robust new baseline for multi-view tracking under unsynchronized conditions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.