acceptodds
Under review as a conference paper at ICLR 2027

CoSlot: Clip-Level Joint Slot Binding for Video Object-Centric Learning

Abstract

Video object discovery requires information from multiple frames while main- taining object identities across frames. In commonly used training setups with per-frame reconstruction, adding context frames also increases the cost of dense de- coding. We propose CoSlot, which jointly binds features from multiple frames with a set of shared slots and combines two reconstruction objectives: shared prototypes reconstruct all visible frames using the binding assignments, while a spatial decoder reconstructs sampled frames from frame-level slot features. This design allows the number of frames used for binding to grow while keeping the number of spatially decoded frames fixed. Experiments show that combining the two reconstruction objectives improves object grouping in videos. In inference comparisons using fixed models, shorter independent chunks mainly reduce cross-frame consistency, and slot passing recovers most of the loss. Without a slot contrastive loss or explicit slot regularization terms, CoSlot achieves video FG-ARI scores of 72.2 and 81.9 on MOVi-C and MOVi-E, respectively. In the evaluated configurations, its cost per optimizer update on precomputed features is lower than that of SlotContrast and SRL. Experiments on YouTube-VIS also show the benefit of passing slots across chunks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.