InteractiveVLN: Proactive Interaction via Uncertainty-Aware Decision for Zero-Shot Vision Language Navigation
Abstract
Most existing Vision-Language Navigation (VLN) methods are evaluated in static environments with single-turn instructions, limiting their ability to generalize to tasks in dynamic, interactive, realistic settings. Moreover, pre-trained navigation policies often struggle to understand complex natural language instructions, tending to overfit to narrow, task-specific directives. To address these challenges, we propose a zero-shot, interaction-enabled VLN agent framework, complementary to task-specific fine-tuning, in which the agent can actively query for task-relevant information during navigation. Unlike traditional model-based approaches, our framework centers on a multimodal long-short memory and an uncertainty-triggered Ask action. This integration enables the agent to anticipate and navigate unseen scenarios when its uncertainty is high. We evaluate our method on the FreeAskWorld and R2R-CE public benchmarks. Without any task-specific fine-tuning, our method achieves the first non-zero success among the zero-shot VLN methods we compare on FreeAskWorld. Notably, the agent autonomously determines when to request guidance: when memory is removed, it seeks assistance more often because uncertainty rises, rather than navigating blindly. These results highlight the potential of our approach for embodied AI and suggest a promising direction for deploying agents in dynamic, collaborative, realistic environments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.