Sequence-Level Post-Training for Visual Tracking
Abstract
Visual trackers are typically trained with frame-level objectives but deployed with sequential inference, where every prediction determines the next search region. This mismatch limits their ability to optimize the behavior encountered at test time. We present Sequence-level Policy Optimization for Tracking (SPOT), a framework for sequence-level post-training. SPOT represents a tracker's native localization output as a stochastic policy, rolls out a group of trajectories under the same prediction-driven cropping mechanism used at inference, and optimizes trajectory-level IoU returns with a critic-free policy gradient. Multiple trajectories provide relative supervision, while frozen-reference KL regularization constrains the updated tracker near its pretrained behavior. Through a unified policy interface, the same framework applies to one-stream, autoregressive, and memory-based trackers without redesigning or modifying their backbone or prediction-head architectures, or changing their test-time inference pipelines. SPOT consistently improves their supervised base trackers across multiple tracking benchmarks. Code will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.