RG-VLN: Benchmarking Underspecified Request Grounding in Vision-Language Navigation
Abstract
Vision-language navigation (VLN) and embodied navigation benchmarks span specified goals, route instructions, target descriptions, and user intentions. In everyday interaction, users may express an underspecified request that combines intent, multiple entities, and partial spatial cues while providing incomplete navigation information. Fulfilling such requests jointly requires target grounding, spatial reasoning, and route planning from the request and online observations. We formulate this complementary setting as Request Grounding VLN. To bridge this research gap, we introduce RG-VLN, a benchmark with over 6,000 episodes across MP3D and HM3D scenes, with both VLM-generated and human-written requests. RG-VLN associates trajectories with detailed route instructions, route-preserving paraphrases, and underspecified requests containing progressively less route information, thereby enabling controlled analysis of navigation specification. Replacing detailed instructions with underspecified requests reduces the success rates of representative VLN methods by 68.5%–74.2%, whereas route-preserving paraphrases cause only minor changes. Performance is also consistently lower for requests with lower navigation-detail density, indicating that reduced navigation specification is a major source of difficulty. Building on these findings, we further propose HiNav, a hierarchical framework that connects request reasoning with navigation execution through closed-loop, scene-grounded textual guidance. This interface resolves request ambiguity from observations and enables existing VLN policies to serve as plug-and-play executors. It achieves the highest success rate across all RG-VLN splits, outperforming the strongest baseline by 7.2%–21.2% relatively. Code and datasets will be publicly released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.