TORUS: Scaling Open-Vocabulary Tracking with Heterogeneous Data and State-Space Models
Abstract
Open-vocabulary multi-object tracking (OVMOT) is short of training data: images with large vocabularies are plentiful, but annotated tracking videos are scarce and temporally sparse. Existing methods either simulate motion from images—which poorly represents real-world deformation and occlusion—or label video by hand, which fails to scale. We introduce TORUS, an online tracker that bypasses this trade-off by unifying both data sources, jointly learning from detection images, disjoint-vocabulary videos, and synthetic pseudo-videos—without using any TAO annotations. TORUS freezes an open-vocabulary detector, adding a selective state-space (Mamba) memory to propagate tracks across frames. To learn jointly from this data mixture, we introduce two key designs: i) a class-agnostic confidence head that decouples track lifecycle management from classification, keeping tracks alive when predicted classes flip; and ii) a spatial-semantic federated loss that masks objectness supervision to avoid background penalization where labels are incomplete. With a ResNet-50 backbone, TORUS achieves 42.5 Base and 42.1 Novel TETA on TAO OVMOT, outperforming the state of the art by +2.9 and +6.8 points. A single set of untuned weights generalizes zero-shot across DanceTrack (57.7 HOTA), YouTube-VIS 2021 (42.9 mAP), and VisDrone-MOT (42.3 TETA)—operating entirely online at 14.7 FPS under 1,203 prompt
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.