Aligned-VLN: Modeling Navigation Progress over Ordered Instructions
Abstract
Many recent Vision-and-Language Navigation (VLN) methods based on vision-language models directly predict actions from natural-language instructions and visual observations. While action supervision specifies what action to execute, it provides limited guidance about which instruction content should guide decisions at different navigation stages. In long-horizon navigation, a locally feasible action may be premature under the instruction order, making action selection dependent on execution progress. We refer to this challenge as instruction-progress mismatch. We propose Aligned-VLN, an action-chunk VLM policy that models navigation progress relative to the ordered instruction. Our key idea is to learn an instruction-conditioned latent alignment over the original instruction tokens. The alignment yields an instruction-relative progress coordinate and an attended instruction representation, which are fused with a learned navigation-state representation into a progress context for action decoding. This context conditions the decoder through layer-wise progress attention and final readout injection. The latent alignment is learned from weak trajectory-derived supervision without token-level execution annotations, with monotonic supervision encouraging progress estimates to follow trajectory order. At each decision step, Aligned-VLN predicts an action chunk with one VLM forward pass, without generated progress descriptions, predefined sub-instructions, or an external progress model. On R2R-CE Val-Unseen, Aligned-VLN achieves 63.5% SR and 56.3% SPL. on RxR-CE, it achieves 59.6% SR and 51.5% SPL. Separate R2R-only controlled experiments show improvements over an action-only policy using the same backbone and action interface.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.