PointWorld: Explicit 3D Representation Calibration for Interactive World Models
Abstract
Existing interactive world models accumulate geometric inconsistencies during long-sequence generation, making it difficult to maintain coherence over thousands of frames. Recent methods mostly provide visual guidance using RGB conditions rendered from 3D representations, but they still suffer from imperfect rendering and reconstruction conflicts that hinder effective geometric conditioning. We find that directly embedding 3D representations can provide geometric guidance while bypassing errors introduced during RGB rendering. Building on this finding, we propose PointWorld that explicitly uses 3D point clouds as geometric memory, rather than relying on rendered RGB proxies, for long-horizon consistency. The core is an explicit 3D representation calibration mechanism that selects relevant points from dynamically reconstructed point clouds, encodes their geometric information, and directly injects the resulting features into the video generation model. To provide reliable 3D conditioning, PointWorld further incorporates a spatiotemporal retrieval strategy into dynamic point cloud reconstruction to mitigate reconstruction conflicts. Extensive experiments demonstrate that PointWorld extends consistent interactive generation, maintaining long-term consistency across sequences of up to 2,000 frames and outperforming recent methods in both quantitative and qualitative comparisons. Our code will be made publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.