AMOR: A Memory-guided Omnidirectional RGB-D Diffusion Model for Geometry-Consistent Next-View Synthesis
Abstract
Visual world models aim to support viewpoint exploration while preserving a coherent environment. Repeatedly extending video clips can compound geometric drift and visual degradation. For static scenes, panoramas cover more space from sparse viewpoints, but each generated view must be internally consistent in appearance and depth, and anchored to prior observations. We present AMOR, a memory-guided paradigm that progressively expands a persistent, explicit 3D scene through joint omnidirectional RGB-D generation at sparse viewpoints. Surface-guided memory readout conditions generation on target-aligned evidence from previously observed surfaces, while conservative updates integrate geometrically compatible predictions into the shared 3D scene. Beyond generating new views, AMOR’s joint modeling of appearance and geometry also enables depth estimation conditioned on observed RGB. Conservative memory updates reduce final-map geometric accuracy error by 21% relative to direct fusion. For novel-view synthesis, AMOR achieves the lowest LPIPS and FID on each of three public benchmarks and the highest average PSNR and SSIM. For depth estimation from observed RGB, it achieves the lowest RMSE and highest δ₁ on TartanGround among the evaluated methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.