Video Foundation Models Bind Objects but Do Not Maintain Them Under Occlusion
Abstract
Video foundation models are widely assumed to acquire object-centric structure, yet whether that structure is causally used, and whether it survives the moment an object stops being visible, has not been tested. Decoding is blind to both: linear readouts recover object identity from frozen features, yet a probe trained to report the hidden object retrieves the occluder instead. We measure object binding inside the model's own prediction machinery: delete object A's early tokens from a frozen predictor's context, then read the per-object degradation of its future predictions. A leakage ratio L summarizes this against endpoints measured in the same pipeline, at K+1 forward passes per clip with no fitted parameters. In a predictive video foundation model (V-JEPA2) and a diffusion one (CogVideoX), objects are strongly bound, and in V-JEPA2 the binding is learned: L = 0.23–0.43 against untrained 1.00 and label-shuffled 1.16, surviving distance- and appearance-matched controls and generalizing to real video. The same models have no object permanence. Under occlusion the fill contrast POperm, hidden object minus occluder, is never significantly positive on our pilot scenes in five video foundation models spanning three training paradigms; where it turns positive at scale (V-JEPA2-H) it is as positive when the object never existed, and V-JEPA2's significant deficits survive subtracting the readout's own measured floor. An object-centric model with per-object recurrent state reaches +0.264; on the 62 clips both cover, the paired gap is +0.292 with the object-centric model higher on 95% of them. The obvious fix works and does not do what it appears to: two epochs of occlusion-completion fine-tuning turn permanence positive (−0.024 → +0.070 over three seeds), and eight lift occluded-object identification from 0.369 to 0.522, yet a paired counterfactual in which the object never existed reproduces 96–108% of the fill gain, identifying the fix as a largely history-blind de-occlusion prior. On this class of feedforward predictors, internal binding does not imply object permanence, and the cheap fix does not install it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.