acceptodds
Under review as a conference paper at ICLR 2027

Which Latent Representations Support Object-Permanent Generation?

Abstract

Video generators often fail to preserve objects that leave the view: occluded ob- jects change identity, fail to return or evolve inconsistently. We ask whether the latent a generator is trained on can natively support object-permanent generation. Our intuition is that a latent encoding the whole scene across cameras and time, rather than one camera’s partial view, can represent objects beyond their visibil- ity. We build PERMAKUBE, a synchronized multi-camera dataset with controlled occlusions and object-level annotations, design PERMATOK, a scene tokenizer trained to render unobserved cameras and times, and train generators on frozen video and scene latents. Under every tokenizer supervision we test, scene la- tents preserve occluded objects more often than video latents despite lower recon- struction fidelity (78.7% versus 47.7% preservation), while accurate motion after occlusion remains challenging for all latents. Cross-view targets add what same- camera targets lack: their encoded scene tokens let a probe trained only on visible objects recover objects hidden from one camera, and their generations reach twice the hide-and-reappear success at similar identity preservation. PERMATOK also adapts to monocular input while retaining strong object-permanent generation, and scene latents are effective diffusion targets built from sparse multi-camera ob- servations. These findings favor scene latents for object-permanent generation and show that identity preservation does not guarantee a readily recoverable hidden- object state, motivating latent probing alongside frame-level evaluation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.