acceptodds
Under review as a conference paper at ICLR 2027

MemWorld: Explore, Edit, and Revisit with Validity-Aware Multimodal Memory

Abstract

Interactive video world models support two complementary interaction modes: users can wander through generated environments using camera actions or direct their evolution through semantic interventions. Existing models typically handle these modes separately. We unify them within one evolving session, enabling users to explore, edit, continue, and revisit. This setting exposes a memory challenge: limited context causes forgetting during long-horizon revisits, while an edit leaves pre-edit observations geometrically useful but visually obsolete. Reusing them as RGB can revert the edit, whereas discarding them loses scene structure. We introduce MemWorld, a memory-augmented video world model with validity-aware multimodal memory. It retrieves nonlocal observations through camera-frustum overlap and uses a version-conditioned RGB–depth policy to preserve current appearance in RGB while retaining obsolete observations as structural depth. Retrieved memory, recent history, and current latents are jointly processed by a camera-conditioned video DiT. We evaluate MemWorld on our proposed editing benchmark of 500 cases covering global appearance editing, continued navigation, and pose-matched revisitation, together with R2M-Bench and MBench. Among the evaluated learned systems, MemWorld achieves the highest R2M-Bench point estimate and the highest Composite point estimate for persistent editing and geometry, while remaining competitive on MBench. These results show that multimodal memory can unify wandering and directing while preserving long-term scene consistency and persistent world updates.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.