ReDual-VLN: Recoverable Dual-System Vision-and-Language Navigation via Language-Conditioned Backtracking
Abstract
Vision-and-language navigation (VLN) aims to enable embodied agents to navigate unseen environments by following natural-language instructions. Long-horizon VLN is vulnerable to route deviations during local execution, which can disrupt the remaining task. Recovery requires selecting an appropriate recovery anchor and maintaining task context for subsequent navigation, placing additional reasoning and memory demands on the VLN model. We propose ReDual-VLN, a dual-system framework that separates high-level reasoning from local navigation execution. A VLM-based Route Supervisor performs route planning, progress tracking, and recovery reasoning, while Hierarchical Context Management (HCM) maintains semantic plans, execution history, and recovery anchors. A Backtracking-Enabled VLN (BE-VLN) then focuses on executing local navigation and recovery instructions. To equip BE-VLN with the required execution capabilities, we construct ReTrace-RxR with paired short-horizon navigation and backtracking data, enabling the model to follow local instructions, predict instruction termination, and perform language-conditioned backtracking. When a deviation occurs, the Route Supervisor selects a previously visited anchor on the intended route as the recovery target and instructs BE-VLN to backtrack to it. After verifying the return, the supervisor directs the executor to resume the remaining instructions, forming a closed loop of deviation detection, backtracking-based recovery, and task resumption. Systematic evaluations in simulation and the real world demonstrate that ReDual-VLN improves long-horizon navigation performance and substantially enhances recovery and task resumption following route deviations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.