The Generator Is the Tracker: Multi-Object Tracking by Painting Persistent Identity Colours
Abstract
Multi-object tracking (MOT) is conventionally decomposed into detection followed by association, with object identity maintained as external state: track buffers, motion models, appearance embeddings. We ask whether a video generator can maintain that state in pixels. We fine-tune a 22B text-to-video diffusion model (LTX-2.3) with a lightweight in-context LoRA to translate an RGB clip into an ID-map clip, a video in which every person is painted a flat, distinct colour that persists over time (same colour, same identity). Long videos are generated as chained windows, each conditioned on the cleaned tail of the previous one; a brief continuation fine-tune teaches the model to extend a given colouring, after which identity flows through the chain with no separately trained detector, no motion model, and no association module. On the DanceTrack test server the system reaches 40.3 HOTA, to our knowledge the first entry whose tracks are read directly off video produced by a pretrained video generator. This is well below today's specialists (66–76 HOTA), but with an error profile no baseline shares: its association score (AssA 44.1) exceeds every tracker of the 2020–22 benchmark suite while detection is the deficit. Controlled comparisons on the validation split show where the identity comes from: windows generated without generation-time state and linked by post-hoc association score 18.2 HOTA against 34.8 for the chain, while re-linking the chain's own windows post hoc changes nothing (34.7), so the colours already carry what overlap association can recover, and motion- and appearance-based associators run on the system's own boxes reach only 22–24 AssA, an ordering that holds again on SportsMOT validation. On 383 mined occlusion events the chained generator re-acquires 42% of the identities it held on both sides of the gap, including gaps longer than its 49-frame context, by keeping the colour alive in pixels through the gap. We release code, checkpoints, and the full pinned experiment log.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.