Learning When Temporal Consistency Should Begin in Video Object-Centric Learning
Abstract
Unsupervised video object-centric learning must both discover new objects and preserve representations of objects already present in the scene. Slot–slot contrastive learning (SSC), a common temporal-consistency objective, favors the latter by encouraging each slot to match its previous-frame state. We show that stronger SSC improves representations of existing objects but progressively degrades representations formed for newly appearing objects, even when idle slot capacity remains available. This reveals an acquisition–maintenance trade-off in how temporal consistency is imposed. We introduce SLOTCOMMIT with learned-onset SSC, which learns a slot-specific onset for temporal consistency to support new-object acquisition without sacrificing the consistency of established representations. This onset is tied to a learned activity gate that determines whether a slot contributes to reconstruction. Upon first activation, the slot becomes committed, and its temporal-consistency constraint begins on the following transition. The first enabled SSC loss also trains the activation decision through commitment, coupling slot recruitment to a representation-dependent consistency cost. On matched MOVi-C models, learned-onset SSC increases acquisition of new objects by idle slots from 27.5% to 50.9% and improves simultaneous entrant–resident representation on separate slots from 25.5% to 33.2%. It also improves overall acquisition and both entrant and resident representation quality. Ablations show that both delaying the onset of SSC and training this onset through commitment contribute to the observed gains. On standard video grouping benchmarks, the complete SLOTCOMMIT model also improves over SlotContrast across MOVi-C, MOVi-E, and YouTube-VIS.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.