Track-OPD: Learning from Future Consequences for Multi-Object Tracking via On-Policy Distillation
Abstract
Online multi-object tracking associates observations using the identity history created by earlier decisions. Learning from the future consequences of these decisions requires comparing alternative associations, yet each alternative can change which tracks survive and which objects their identities denote. We propose Track-OPD, an on-policy distillation method that makes these consequences comparable under fixed identity bindings. Alternative associations continue from the same tracker state under shared future observations and a frozen policy. Their outcomes are evaluated on the same annotated objects and frames, retaining missing and incorrect predictions in the comparison. The resulting future identity preferences are aligned with current association classes to construct a local teacher. This teacher redistributes probability between the ground-truth reference and verified poorer alternatives while preserving their total probability. On DanceTrack and SportsMOT, Track-OPD improves HOTA over continued training by 1.03–1.36 points across the evaluated tracker adaptations. At 4,000 updates, its target construction gains 0.543 HOTA and 1.184 AssA over unit ground-truth preferences on the same selected records. These results support learning current associations from the identity history they produce while retaining the original online inference structure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.