SLES: Self-Localized Entity Supervision for Long-Horizon Dynamic Memory in Video World Models
Abstract
Long-horizon simulation of a persistent world requires maintaining the state of dynamic entities even while they lie outside the field of view, recovering each with its identity and evolving state intact upon re-entry; we refer to this capability as long-horizon dynamic memory. Despite rapid advances in generation quality and controllability, existing models frequently fail to recover entities that have left the field of view, or do so with a different identity or implausible dynamics. We study this problem under full-context bidirectional generation, where the complete observation history remains accessible, and find that failures persist, indicating that information utilization, rather than access, is the bottleneck. Probing self-attention reveals an observability–controllability gap: entity retrieval is encoded in reproducible attention structure, yet directly steering this routing does not reliably improve generation and can degrade identity consistency. We therefore use attention only to localize intervention. In this paper, we propose Self-Localized Entity Supervision (SLES), which derives a lightweight spatial entity support from text prior and reuses it to up-weight entity-aware supervision during training and to localize a Langevin-inspired overshooting sampler that concentrates stochastic correction on the entity region during inference, requiring no external memory, architectural modification, or target-side masks. We further introduce a phase-aware protocol separating reappearance, subject consistency, and temporal recovery. Across backbones and ablations, self-derived support approaches ground-truth mask supervision for identity recovery, while localized overshooting improves post-re-entry dynamics with limited collateral disturbance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.