Geometry-Conditioned Latent World Models with Recurrent Adaptation in Dynamic Environments
Abstract
Autonomous navigation frequently relies on pre-mapped 3D spatial priors, but real-world deployments are inherently dynamic, resulting in inconsistencies between static maps and observations. Latent world models offer a promising, lightweight solution operating solely on RGB video, but current architectures lack the explicit spatial understanding and persistent memory to plan effectively in dynamical environments. To bridge this gap, we introduce GeoMamba-WM, a geometry-grounded recurrent world model that encodes a pre-mapped semantic 3D point cloud and utilizes a Mamba-based State Space Model (SSM) to dynamically correct a persistent hidden state using live visual observations. Prediction and planning are performed within a fused latent space derived via cross-attention between the observation and corrected geometry prior and stabilized by an auxiliary objective that encourages information captured by the SSM state to persist throughout the predictive architecture. To systematically evaluate this, we introduce GOI-NAV, a 3D navigation benchmark that features geometry-observation incongruences such as displaced and unseen obstacles, deformable surfaces, and visual noise. Empirical evaluations demonstrate that GeoMamba-WM significantly outperforms baseline latent world models across all incongruence axes, establishing a new standard for offline, reward-free predictive planning. Ultimately, our results show that latent world models are highly viable for planning in dynamic, pre-mapped 3D environments using depth-less RGB input, demonstrating that cross-modal inconsistencies can be resolved entirely in latent space without explicit mapping or online reconstruction objectives.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.