QTrack: Learning to Use Temporal Evidence for Reference-Grounded Video Tracking
Abstract
Category-level multi-object trackers estimate trajectories for all detected instances, whereas some applications require propagating only targets specified by a user. We study offline, reference-grounded video tracking: given an ordered video clip, a natural-language query, and the reference-frame boxes of the specified targets, a model predicts their subsequent trajectories. We introduce QTrack, a vision-language policy trained with structured trajectory rewards and Temporal Perception-Aware Policy Optimization (TAPO). TAPO contrasts the likelihood of a generated response under the observed clip and a reference-frame repetition, encouraging the output to depend on visual evidence beyond the reference frame. We also introduce RMOT26, a controlled diagnostic benchmark constructed from existing tracking datasets. Under the reported protocol, QTrack obtains a motion-consistency score of 0.30, MOTP of 0.75, center-location error of 44.61 pixels, and normalized distance error of 0.39; its motion-consistency score ties the strongest comparison method. Component ablations suggest complementary effects from the motion reward and temporal contrast, while an exploratory four-clip, 50-frame evaluation examines transfer beyond the primary short-clip setting. The results support further study of temporal-evidence learning in trajectory-generating vision-language models. Code and data are available at https://anonymous.4open.science/r/QTrack-5548.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.