Anchors Guided Latent World Model for Long-Horizon Vision-and-Language Navigation
Abstract
Vision-and-Language Navigation (VLN) remains challenging, requiring agents to ground language instructions into actions under partial observability. In long-horizon VLN tasks, accumulated trajectory errors during planning lead to progressive deviation from the goal, degrading navigation efficiency. Recent methods introduce world models to anticipate future navigation states and improve candidate action evaluation, which can be broadly categorized into two categories. One line predicts the next state for local action refinement, while the other performs multi-step prediction for long-horizon planning. However, both classes of methods lack an explicit representation of global structure, making them susceptible to inefficient exploration in semantically ambiguous environments. To address this limitation, we propose AnchorWorld, which introduces a set of sparse prior anchors, explicitly supervised by future navigation states, to guide latent-space action generation within the world model, thereby enhancing global structural awareness and improving navigation efficiency. Extensive experiments on the R2R and R4R datasets demonstrate the effectiveness of the proposed method. The advantages are most pronounced in long-horizon navigation, demonstrating improved efficiency and robustness in long-term decision-making.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.