acceptodds
Under review as a conference paper at ICLR 2027

MC-World: Memory-Guided Video Diffusion with Collision-Aware Exploration for Scene-Consistent 3D World Generation

Abstract

Diffusion-based 3D scene generation has achieved remarkable progress. However, existing methods still face two critical challenges: (1) They struggle to remember the historical scenes and their spatial relationships in long-range video generation, making it difficult to maintain long-term 3D scene consistency; (2) They are prone to trajectory collisions with objects due to the lack of evasion capabilities for objects, leading to unreasonable geometric structures and inconsistent temporal sequences, thereby further limiting realistic and smooth scene exploration. To address these challenges, we propose MC-World, a novel memory-guided and collision-aware video diffusion framework for explorable 3D world generation. Specifically, MC-World consists of two core designs: Spatio-Temporal Memory Retrieval and Collision-Aware Exploration Control. The former retrieves temporally coherent and spatially relevant historical frames from a global memory cache, providing long-range contextual guidance for consistent multi-view generation. Collaboratively, the latter leverages collision detection strategy to adaptively truncate high-risk trajectory and switch generation directions, thus ensuring geometrically reasonable and temporally consistent scene content. Benefiting from the above designs, our method enables consistent 3D long-range scene generation and supports unrestricted realistic world exploration. Extensive experiments show that MC-World significantly superior to existing methods in various scenes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.