SeaMem: Seen-Aware Implicit Memory for Long-Horizon Camera-Controlled World Generation
Abstract
Long-horizon world generation aims to enable controllable exploration of a scene beyond initial observations by generating views along user-specified camera trajectories. As the viewpoint changes, each target view may contain both previously seen content and newly revealed regions. The former must remain consistent with historical observations, while the latter requires plausible synthesis guided by learned world priors. The key challenge is therefore how to represent and use historical observations to preserve previously seen content while allowing plausible synthesis in newly revealed regions. Existing methods typically rely on either explicit 3D memory or implicit context memory. However, 3D reconstruction can accumulate persistent geometric errors when integrating imperfect generated views. In contrast, implicit context memory avoids this reconstruction error, but lacks awareness of seen and unseen regions. Consequently, previously observed content may be unnecessarily resynthesized, introducing hallucinations and degrading long-horizon consistency. To address these issues, we introduce SeaMem, a seen-aware implicit memory for long-horizon camera-controlled world generation. SeaMem couples implicit recovery of historical content with explicit identification of where the recovered content should constrain generation. Specifically, SeaMem queries compact historical representations with the target camera to recover content and predicts an unseen-region mask from attention statistics. Recovered content anchors seen regions, while a pretrained novel view synthesis (NVS) backbone synthesizes unseen regions. Within seen regions, a detail-deficient mask relaxes anchoring where historical detail is insufficient, enabling refinement under historical constraints. This design extends short-range NVS backbones to long-horizon world generation without requiring additional backbone fine-tuning. Extensive experiments demonstrate that SeaMem effectively extends NVS backbones to long-horizon world generation while achieving robust revisit consistency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.