Evidence-Gated State Anchoring for Persistent Video World Models
Abstract
Pretrained video world models change regions that nothing has touched, and they fail to restore a view after the camera looks away and comes back. Both behaviors come from the same gap. Self-attention can read any token in the window, while the training loss never asks whether a location was licensed to change. Recent memory-augmented world models decide which history stays readable. They leave the local write unconstrained. Evidence-Gated State Anchoring (EGSA) is a frozen-backbone retrofit that turns correspondence and propagated-change evidence into write authorization. Where that evidence is absent, a residual hold mixes the clean-state prediction toward a warped copy of the model's own previous estimate, inside every denoising step. Training touches fewer than 0.4 percent of parameters. A warmup multiplier and zeroed LoRA output projections make the retrofit identical to the base model at step zero, and the evidence detector sees only the noised model input. A no-bias ablation matches the full model to within 0.5 percent on both backbones, which places the gain on the anchoring pathway. On a 0.5B latent flow-matching model and the 0.46B pixel-space DFoT RealEstate10K model, EGSA improves real-capture palindromic revisits against a matched LoRA twin. On DFoT, look-away error drops by 67 percent and near-origin return error by 12 to 14 percent (86 to 93 percent of held-out scenes, scene-level p < 10^-4), across training seeds, rollout noise seeds, and a monotone inference dose. On the flow-matching model, all three training seeds lower return error on 128 stronger-departure held-out scenes, by up to 12 percent. Revisit targets are captured frames, fidelity is reported together with a view-change tracking ratio, and all revisit scores are computed in decoded pixels.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.