GLAM: TRAINING A LATENT WORLD MODEL OVER GLOBAL SPATIOTEMPORAL MEMORY FOR ACTIVE EXPLORATION AND NAVIGATION
Abstract
Active exploration and semantic navigation require an embodied agent to build memory from partial observations, predict how the evolution of observed spatial memory may support future motion, and convert that prediction into actionable plans. We present GLAM, a goal-conditioned latent world model trained over global spatiotemporal memory, and GLAM NAV, the complete navigation system built around it. Given historical map tokens, a navigation goal, and the current robot pose, GLAM jointly predicts future map representations and robot-centric waypoint latents, allowing future spatial context and navigation intent to be inferred in a shared representation space. The model follows a JEPA-like latent prediction paradigm, operates directly on map-level latent tokens rather than RGB reconstruction, and uses a pretrained waypoint encoder–decoder to supervise and decode navigation plans within GLAM NAV. Training data are collected in Habitat over HM3D v0.2 by sampling oracle-guided ObjectNav trajectories and their map transitions. On the full HM3D-Semantics v0.2 ObjectNav benchmark, GLAM NAV achieves a state-of-the-art success rate of 82.30% with 34.00% success weighted by path length.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.