GroundingVLN: Reasoning and Acting with Grounding for Vision-Language Navigation
Abstract
Although vision-language models (VLMs) possess strong visual understanding and reasoning capabilities, vision-and-language navigation (VLN) agents built upon them still struggle to bridge semantic reasoning and spatial execution. This disconnection stems from two coupled limitations: intermediate reasoning is not explicitly grounded in visual evidence, and high-level decisions lack precise spatial goals to guide low-level motion. Inspired by human navigation, which naturally links cognition and locomotion through landmarks and spatial goals, we propose **GroundingVLN**, a framework that uses visual grounding as a shared interface between reasoning and action. GroundingVLN first reasons with grounding by anchoring task-relevant visual evidence to precise image locations throughout structured reasoning. It then acts through grounding by predicting a progress-aligned pixel goal that a geometric planner translates into primitive actions. To learn these capabilities, we construct **GroundingCOTVLN-188K**, a dataset of temporally aligned grounded reasoning traces, and introduce **Grounded and Execution-Aware Reinforcement Learning (GEAR)**, which aligns grounded reasoning and spatial decisions with downstream execution. Experiments demonstrate that GroundingVLN achieves state-of-the-art performance (**69.9%** SR on R2R-CE and **75.1%** SR on RxR-CE) with superior data efficiency, requiring only **0.9%** of the training data used by the strongest baseline. It also generalizes strongly across datasets, attaining **59.9%** SR on RxR-CE when trained solely on R2R, a gain of **20.1%** over the strongest baseline. The full code, dataset, and model will be released after review.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.