acceptodds
Under review as a conference paper at ICLR 2027

Learning Composable Transition Representations for Visual Navigation

Abstract

What latent representation should an image-goal navigation policy use? Direct visual policies are efficient, but their representations are often learned primarily through action supervision and may fail to capture the transition structure underlying viewpoint changes. We propose NavCompose, a self-supervised pretraining paradigm that learns structured navigation latents from unlabeled visual trajectories. NavCompose encodes relative visual transitions and explicitly regularizes them to be composable across consecutive observations and continuous along local trajectories. The learned representation is renderable and supports interpolation between observed viewpoints as well as short-horizon extrapolation beyond them, providing a compact model of local visual transitions. For downstream image-goal navigation, we freeze the pre-trained backbone and train only a lightweight waypoint prediction head. Across multiple navigation datasets, NavCompose achieves state-of-the-art waypoint prediction performance while maintaining strong accuracy with substantially fewer labeled trajectories. Its lightweight downstream architecture further enables fast inference, outperforming existing navigation models in runtime efficiency. Code will be released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.