Reading Point Tracks from Frozen Video Diffusion Transformers
Abstract
We show that readout design substantially improves point tracking from frozen video diffusion transformers, without training a tracking network. Our readout combines confidence-weighted attention heads, bilinear source sampling, local position decoding and visibility estimation from one transformer block. All settings are selected on 20 labelled synthetic clips and transferred unchanged to real videos. CogVideoX-2B reaches 5.21 average Jaccard (AJ) on TAP-Vid DAVIS and 44.00 on Kinetics. Evaluating released HeFT on the same videos, queries and checkpoints with the official TAP-Vid first-query metric establishes gains on both benchmarks for CogVideoX and Wan, including +18.62 and +12.77 AJ on Kinetics. Controlled ablations on both backbones show why localisation improves: bilinear source sampling and local decoding reinforce each other, and most of the local-decoding gain comes from predicting positions between token centres. On CogVideoX, confidence weighting further improves head aggregation, whereas tested block combinations give no clear gain on the synthetic selection set. Across six checkpoints from three generator families, larger models do not consistently improve tracking under this readout. Our readout improves zero-shot correspondence estimation while keeping the video generator frozen.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.