acceptodds
Under review as a conference paper at ICLR 2027

TrackDiT: 3D Tracking from 2D Tracks and 3D reconstruction

Abstract

Recent 3D point trackers adopt architectures that jointly learn appearance matching, 3D reconstruction and temporal reasoning in a single yet complex end-to-end model. By evaluating a simple baseline that lifts 2D tracks into 3D using an off-the-shelf 2D tracker and 3D reconstructor, we derive two insights: (i) such a baseline outperforms specialized 3D tracking models on always-visible tracks, and (ii) when tracks are partially occluded, specialized models surpass such a baseline. This observation motivates a new paradigm: rather than training one network for everything, we first obtain precise visible track fragments from foundation models, then treat the remaining occlusion gaps as a track completion problem, addressed by a flow-matching model that generates the correction to the lifted trajectory. Because a hidden point’s motion is genuinely ambiguous, a generative formulation is a better fit than regression, and operating on track-frame tokens lets the model complete a gap both from the track’s own motion and from the co-visible points around it. The resulting framework is modular, simple to train, and automatically improves as its foundation-model components advance. TrackDiT achieves state-of-the-art 3D tracking results on ADT, PStudio, and SynthVerse. In particular, on occluded points it improves over the lifting baseline by 12.2 APD3D points and surpasses specialized models, while also improving accuracy on visible points.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.