iSEE: Object Permanence Through Self-Supervision
Abstract
Object permanence, keeping track of an object’s identity and position while it is occluded, is central to video representations that track, predict and plan. Trackers that achieve it learn from boxes, track identities and visibility labels. On the other hand, self-supervised object-centric methods discover objects without labels: through slot attention, it represents a video as slots that bind to objects and follow them across frames. However, these slots are lost under occlusion, making the desired permanence impossible. Reasoning permanence is a hard problem because it requires to detect when an object becomes occluded, re-identify when object reappears, and keep the object’s hidden position continuous, using reapperance as the only learning cue. To address this, we propose iSEE, a novel framework that offers all three aforementioned requirements, without any labels whatsoever. We built iSEE using the following three proposed components: (i) Object evidence modelling: a slot’s attention, compared with its own past, reveals when its object is hidden. (ii) Appearance-position separation: two slot streams let the appearance be held for re-identification while the position keeps changing. (iii) Permanence from reappearance: a walker follows the hidden object’s position, trained only on where the object reappears. On LA-CATER static, iSEE returns a reappearing object to its own slot after 86 % of occlusions, against 32 % for SlotContrast (Manasyan et al., 2025), and localises it while hidden within 4.1 mAP of the label-trained SoTA RAM (Tokmakov et al., 2022). The two streams also allow downstream planning, with the position stream as the action of a world model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.