acceptodds
Under review as a conference paper at ICLR 2027

Reading Point Tracks from Frozen Video Diffusion Transformers

Abstract

We show that readout design substantially improves point tracking from frozen video diffusion transformers, without training a tracking network. Our readout combines confidence-weighted attention heads, bilinear source sampling, local position decoding and visibility estimation from one transformer block. All settings are selected on 20 labelled synthetic clips and transferred unchanged to real videos. CogVideoX-2B reaches 5.21 average Jaccard (AJ) on TAP-Vid DAVIS and 44.00 on Kinetics. Evaluating released HeFT on the same videos, queries and checkpoints with the official TAP-Vid first-query metric establishes gains on both benchmarks for CogVideoX and Wan, including +18.62 and +12.77 AJ on Kinetics. Controlled ablations on both backbones show why localisation improves: bilinear source sampling and local decoding reinforce each other, and most of the local-decoding gain comes from predicting positions between token centres. On CogVideoX, confidence weighting further improves head aggregation, whereas tested block combinations give no clear gain on the synthetic selection set. Across six checkpoints from three generator families, larger models do not consistently improve tracking under this readout. Our readout improves zero-shot correspondence estimation while keeping the video generator frozen.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.