acceptodds
Under review as a conference paper at ICLR 2027

QTrack: Learning to Use Temporal Evidence for Reference-Grounded Video Tracking

Abstract

Category-level multi-object trackers estimate trajectories for all detected instances, whereas some applications require propagating only targets specified by a user. We study offline, reference-grounded video tracking: given an ordered video clip, a natural-language query, and the reference-frame boxes of the specified targets, a model predicts their subsequent trajectories. We introduce QTrack, a vision-language policy trained with structured trajectory rewards and Temporal Perception-Aware Policy Optimization (TAPO). TAPO contrasts the likelihood of a generated response under the observed clip and a reference-frame repetition, encouraging the output to depend on visual evidence beyond the reference frame. We also introduce RMOT26, a controlled diagnostic benchmark constructed from existing tracking datasets. Under the reported protocol, QTrack obtains a motion-consistency score of 0.30, MOTP of 0.75, center-location error of 44.61 pixels, and normalized distance error of 0.39; its motion-consistency score ties the strongest comparison method. Component ablations suggest complementary effects from the motion reward and temporal contrast, while an exploratory four-clip, 50-frame evaluation examines transfer beyond the primary short-clip setting. The results support further study of temporal-evidence learning in trajectory-generating vision-language models. Code and data are available at https://anonymous.4open.science/r/QTrack-5548.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.