Learning Composable Transition Representations for Visual Navigation
Abstract
What latent representation should an image-goal navigation policy use? Direct visual policies are efficient, but their representations are often learned primarily through action supervision and may fail to capture the transition structure underlying viewpoint changes. We propose NavCompose, a self-supervised pretraining paradigm that learns structured navigation latents from unlabeled visual trajectories. NavCompose encodes relative visual transitions and explicitly regularizes them to be composable across consecutive observations and continuous along local trajectories. The learned representation is renderable and supports interpolation between observed viewpoints as well as short-horizon extrapolation beyond them, providing a compact model of local visual transitions. For downstream image-goal navigation, we freeze the pre-trained backbone and train only a lightweight waypoint prediction head. Across multiple navigation datasets, NavCompose achieves state-of-the-art waypoint prediction performance while maintaining strong accuracy with substantially fewer labeled trajectories. Its lightweight downstream architecture further enables fast inference, outperforming existing navigation models in runtime efficiency. Code will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.