Identity without Propagation or Matching: Parallel Object-Centric Video Learning
Abstract
Video object-centric learning decomposes a video into a small set of object slots. The goal is that each slot binds one object and keeps that binding in every frame. The methods that lead the video benchmarks secure that binding by propagation, initializing each frame from the previous frame's slots through a learned transition model. We present CoSlot, a fully parallel video object-centric model that binds an entire clip in a joint pass, one slot-attention competition over all of its patches, and then refines every frame independently with the same module and no explicit propagation mechanism. It is trained with weak supervision: a differentiable agreement term in the loss compares the slot assignment, pooled over the whole clip, with imperfect reference masks from an unsupervised model or a foundation segmenter; no ground-truth mask is used anywhere. The identity it obtains is competitive: against three published baselines, CoSlot leads on clip-level identity on two of three benchmarks, by to Clip FG-ARI on YouTube-VIS, and matches the strongest on the third, at comparable mask quality. We provide a theory that explains this empirical performance. The output of the joint pass defines, for each clip, convex cones in the projected feature space, one per slot and the same in every frame; an object whose patches stay in one cone keeps one slot in every frame, because every frame is partitioned by the same cones and no information passes between frames. Empirically, the cones the trained joint pass produces predict which slot owns an object significantly better than those of an untrained model.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.